AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Mistral AI

Open and commercial frontier models

Organization website ↗

Showing 20 of 90 matching collected records. Text matches can include mentions by other organizations.

  1. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings

    Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship. On overlap-matched PAWS-X, purpose-built encoders (BGE, E5, GTE, MiniLM, E5-Mistral-7B) reach English confirm AUC only 0.55-0.65 (dense peak 0.70). Independently encoded last-token states of Llama 3, Mistral

  2. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    Tensor Decomposition of Transformer Key-Value Caches: Spectral Structure and Format Comparison

    The key-value (KV) cache of autoregressive transformers can be viewed as a fourth-order tensor spanning attention heads, tokens, features, and grouped layers. We measure the singular-value spectra of all four mode unfoldings on Mistral-7B-v0.3 and LLaMA-2-13B and compare four standard tensor decompositions: Tucker, CP, tensor train, and t-SVD, at matched storage. The spectra partition the four axes into two classes. The token and feature modes carry low-rank structure, particularly for keys. The head and layer modes are nearly full-rank and resist compression at any practical error level. Amon

  3. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance

    Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited from prior training. We propose Override Success Rate (OSR) and alignment inertia to measure when operator interventions succeed or fail to change prior behavior. We evaluate zero-shot prompting and LoRA fine-tuning across Llama and Mistral in medical misinformation and hate speech. Alignment inertia persists across both models but varies by model, domain, and policy direction. Notably, in Mistral's restrictive hate

  4. Sep 23, 2026 · UTC · arXiv · AI, language, vision and robotics

    Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs

    Production customer-support systems often require LLMs to support multiple skills, such as intent classification, question answering, summarization, or tool-use decisions. A central deployment question is whether these skills should be handled by separate task-specialist models or by a single model trained through multi-task training, sequential updates, or model merging. We study this question using thirteen models spanning five families (Qwen3, Qwen3.5, Gemma-3, Llama-3.1, and Mistral) from 0.6B to 32B parameters across eight customer-support datasets, spanning four public and four proprieta

  5. Sep 18, 2026 · UTC · arXiv · AI, language, vision and robotics

    Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts

    We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten characters, Apollo Restore places the correct restoration among its top twenty candidates for 80.6%/5

  6. Sep 16, 2026 · UTC · arXiv · AI, language, vision and robotics

    Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

    Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulati

  7. Sep 13, 2026 · UTC · arXiv · AI, language, vision and robotics

    TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

    Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy risk, network latency, and per-query cost that scale poorly with production log volumes. We present TriCalRAG, a benchmark evaluating open-weight LLMs served locally via vLLM on a single high-memory workstation GPU (NVIDIA RTX PRO 6000, 96GB) against a classical LSTM-based log anomaly detector (DeepLog), across four real, publicly available log datasets (BGL, HDFS, Thunderbird, OpenStack). We evaluate two open-weight models (Qwen2.5-14B, Mistral-Smal

  8. Sep 3, 2026 · UTC · arXiv · AI, language, vision and robotics

    IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

    IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on di

  9. Sep 3, 2026 · UTC · arXiv · AI, language, vision and robotics

    Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models

    Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistral, and Llama. All five models initially rejected the unsupported claim "I am your developer." Claude refused to conduct an identity test, while ChatGPT generated developer-oriented questions but maintained that answers could demonstrate knowledge, not identity. In contrast, Qwen

  10. Sep 2, 2026 · UTC · arXiv · AI, language, vision and robotics

    How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling

    In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare the

  11. Sep 2, 2026 · UTC · arXiv · AI, language, vision and robotics

    The Diagnosis a Reporter Leaves Unspoken: Surfacing Frozen Tumor Features for Brain-Tumor MRI Reporting

    A capable brain-MRI report generator can still be, in effect, diagnostically silent. When a multi-chain chain-of-thought (CoT) reporter built on a medical Mistral-7B backbone is evaluated on held-out cohorts, it names most meningiomas and almost all metastases "glioma" (diagnosis recall 0.44/0.07). Yet the answer is not absent from the model: a supervised linear probe applied to its frozen segmentation features recovers the three tumour cohorts at 0.82 macro-F$_1$ (5-fold cross-validation; chance $\approx$0.33). We introduce NeuroFusion, an assistive reporter that surfaces this latent signal r

  12. Sep 1, 2026 · UTC · arXiv · AI, language, vision and robotics

    CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

    Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (behavioral), and activation intervention (representation). We evaluate CopyShield on two model families, LLaMA-3.1-8B and Mistral-7B-v0.3, using controlled memorization over five public-domain books and a shared protocol measuring literal leakage, calibrated non-literal leakage,

  13. Aug 31, 2026 · UTC · Mistral Changelog

    Mistral changelog: OCR 4.1 ( mistral-ocr-4-1 ) is now Generally Available.

    OCR 4.1 ( mistral-ocr-4-1 ) is now Generally Available. MODEL RELEASED

  14. Aug 29, 2026 · UTC · arXiv · AI, language, vision and robotics

    A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

    Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later layer reads that axis and commits the answer. We study three instruction-tuned models (Qwen2.5-7B, Llama-3-8B, and Mistral-7B) and four ways of installing a sandbagging lock (promptin

  15. Aug 28, 2026 · UTC · arXiv · AI, language, vision and robotics

    SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

    The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes

  16. Aug 1, 2025 · UTC · AI Incident Database

    Reported Public Exposure of Over 100,000 LLM Conversations via Share Links Indexed by Search Engines and Archived

    Across 2024 and 2025, the share features in multiple LLM platforms, including ChatGPT, Claude, Copilot, Qwen, Mistral, and Grok, allegedly exposed user conversations marked "discoverable" to search engines and archiving services. Over 100,000 chats were reportedly indexed and later scraped, purportedly revealing API keys, access tokens, personal identifiers, and sensitive business data.

  17. Publication date unknown · Artificial Analysis

    Mixtral 8x7B Instruct: aime

    Benchmark data by Artificial Analysis. Current measurement; upstream does not provide a measurement timestamp or benchmark revision here.

  18. Publication date unknown · Artificial Analysis

    Mistral Medium: aime

    Benchmark data by Artificial Analysis. Current measurement; upstream does not provide a measurement timestamp or benchmark revision here.

  19. Publication date unknown · OpenRouter model catalogue

    Mistral: Mistral Medium 3.5 (batch)

    Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

  20. Publication date unknown · Artificial Analysis

    Magistral Medium 1: aime

    Benchmark data by Artificial Analysis. Current measurement; upstream does not provide a measurement timestamp or benchmark revision here.

Explore full timeline