Artificial Analysis’ cover photo
Artificial Analysis

Artificial Analysis

Technology, Information and Internet

Newark, Delaware 26,577 followers

Independent analysis of AI: Understand the AI landscape and analyze AI technologies http://artificialanalysis.com/

About us

Leading independent analysis of AI. Backed by Nat Friedman, Daniel Gross and Andrew Ng.

Website
https://artificialanalysis.ai
Industry
Technology, Information and Internet
Company size
11-50 employees
Headquarters
Newark, Delaware
Type
Privately Held

Locations

Employees at Artificial Analysis

Updates

  • Xiaomi’s MiMo V2.5 Pro has landed at 54 in the Artificial Analysis Intelligence Index, tied with Moonshot’s Kimi K2.6 - the current top open weights model. MiMo V2.5 Pro’s weights are expected to be released soon, which would make MiMo V2.5 Pro the first equal open weights model - slightly ahead of DeepSeek V4 Pro Xiaomi's MiMo V2.5 Pro shows an impressive improvement over MiMo V2 Pro (49), the previous generation of Xiaomi's flagship model family, which was released just over a month ago on March 19, 2026. Key takeaways: ➤ MiMo V2.5 Pro is on the pareto frontier of our Intelligence Index vs Cost to Run Intelligence Index chart. It was slightly cheaper to run than GLM-5.1, and slightly more intelligent. It was significantly cheaper to run than Kimi K2.6, driven by using just over half the number of output tokens. ➤ MiMo V2.5 Pro will be the leading open weights model in GDPval-AA, our agentic real-world work tasks benchmark. It scores 1578, ahead of DeepSeek V4 Pro (1554), GLM-5.1 (1535), MiniMax-M2.7 (1514), and Kimi K2.6 (1484). ➤ It makes progress in reasoning and instruction following. The model scores 34% on HLE (+6% from MiMo V2.0) and 80% on IFBench (+11% from MiMo V2.0). However, compared to the previous generation, there is a small regression in CritPt (5% to 4%). ➤ MiMo V2.5 Pro's token efficiency remains competitive against peers in a similar intelligence tier, using ~92M output tokens for the Intelligence Index. This is more efficient than Kimi K2.6 (~170M) and GLM 5.1 (~110M). However, it does use 19% more than the previous generation model, MiMo V2 Pro (77M). ➤ Priced at $1.00/$3.00 per million input/output tokens on Xiaomi’s first-party API, MiMo V2.5 Pro is relatively cost-efficient for its intelligence tier. It costs only $462 to run the Artificial Analysis Intelligence Index, compared to $948 for Kimi K2.6 and $544 for GLM 5.1. ➤ MiMo V2.5 Pro scores 4 on the AA-Omniscience Index, a proprietary Artificial Analysis evaluation that measures factual accuracy and hallucination. This is a slight regression from MiMo V2 Pro (5), though both models still trail proprietary frontier models. MiMo V2.5 Pro demonstrates a relatively low hallucination rate (25%) but also low accuracy (23%). Additional model details: ➤ Context window: 1M tokens ➤ Parameters: 1T total, 42B active ➤ License: Xiaomi has publicly announced that weights are to be released soon. The model will show on Artificial Analysis as a ‘proprietary’ until the weights are released ➤ Release date: April 22, 2026 ➤ Availability: MiMo V2.5 Pro is available via Xiaomi's first-party API

    • No alternative text description for this image
  • DeepSeek V4 Pro is the top open weights model on GDPval-AA, our agentic real-world work tasks evaluation DeepSeek AI has released V4 Pro (1.6T total / 49B active) and V4 Flash (284B total / 13B active). V4 is DeepSeek's first new size since V3, with all intermediate models (V3.1, V3.2, R1, R1 0528) sharing the V3 family's 685B total / 37B active parameter MoE design. V4 Pro is also the largest open weights model released to date, surpassing Kimi K2.6 (1T total / 32B active) in both total and active parameter counts. V4 Pro is released mostly in FP4 precision, putting total model size at ~865GB, comparable to Kimi K2.6 (INT4, ~500GB). GLM-5.1 is BF16 (~1.49TB) natively and typically served in FP8 or FP4. Both models are hybrid thinking/non-thinking, and we tested the reasoning variants at Max Effort and High Effort.

  • GPT Image 2 (high) debuts at #1 on our Text to Image Leaderboard, surpassing Nano Banana 2, FLUX.2 [max], and Seedream 4.0 in the Artificial Analysis Image Arena. OpenAI's latest image model is a leap forward in prompt adherence, photorealism, and text rendering. We see the biggest difference in our most complex prompts, especially where no model to date has been able to follow the instructions. On our Image Editing leaderboard, GPT Image 2 appears to be much less of a leap forward, landing approximately in line with GPT Image 1.5. Our Image Editing leaderboard tests prompts that make changes to a single input image (eg. change text, remove an object, add a person). GPT Image 2 (high) is priced at $211 per 1k images via API, positioning it at a higher price point than Google’s Nano Banana 2 ($67 per 1k images). It is available via OpenAI’s developer API and in ChatGPT.

    • No alternative text description for this image
  • Ant Group's Ling 2.6 Flash scores 26 on the Artificial Analysis Intelligence Index, a 10-point jump from Ling-flash-2.0. It is one of few recent open weights releases focused on non-reasoning capabilities and focuses on a reasonable cost to intelligence ratio. Ling 2.6 Flash is a non-reasoning model from Ant Group's InclusionAI lab. Ant Group's model family comprises three series: Ling (non-reasoning), Ring (reasoning), and Ming (multimodal). Ling-flash-2.0 was the previous flash-tier non-reasoning model. Ling 2.6 Flash is expected to be open weights shortly after release, but as of today the weights have not been released on Hugging Face. Key takeaways: ➤ At 104B total parameters with 7.4B active parameters, Ling 2.6 Flash (26) sits in intelligence near GPT-5.4 nano (Non-Reasoning, 24) and Gemma 4 26B A4B (Non-reasoning, 27), both models with comparable active parameter counts. However, at 18 points behind GLM-5.1 (Non-reasoning, 44), there remains a gap to frontier non-reasoning open weights models ➤ Ling 2.6 Flash is comparatively token efficient, using ~15M output tokens to run the Intelligence Index. This is comparable to Gemma 4 26B A4B (~14M) but a fraction of Qwen3.5 9B (~78M). Compared to models in the similar intelligence tier, Ling 2.6 Flash represents a reasonable efficiency tradeoff, which has positive effects on cost when deployed on larger workloads. At a price of $0.1 / million input tokens and $0.3 / million output tokens, Ling 2.6 Flash costs only ~$23 to run the full Artificial Analysis Intelligence Index. ➤ Gains from Ling-flash-2.0 were driven mostly by improvements agentic capabilities and instruction following. τ²-Bench jumped from 21% to 86% (+65 points), IFBench from 34% to 57% (+23 points), and GDPval-AA Elo from 425 to 783 (+84%). Conversely, GPQA Diamond fell from 66% to 59% (-6 points) and SciCode from 29% to 27% (-2 points). ➤ AA-Omniscience performance is at -66 with 15% accuracy and 96% hallucination rate. This is consistent with the model's small 7.4B active parameter count. Knowledge recall benefits from larger parameter counts, and sub-10B active-parameter models systematically underperform on this metric. Additional model details: ➤ Architecture: MoE, 104B total parameters, 7.4B active parameters ➤ Context window: 262K tokens (doubled from 128K for Ling-flash-2.0) ➤ Pricing: $0.10 / $0.30 per 1M input/output tokens (via Novita API) ➤ License: Weights not yet released ➤ Availability: Third party API through Novita

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands 4th on the Artificial Analysis Intelligence Index (54) behind only Anthropic, Google, and OpenAI (all 57) Key takeaways: ➤ Increase in performance on agentic tasks: Kimi K2.6 achieves an Elo of 1520 on our GDPval-AA evaluation, which is a marked improvement over Kimi K2.5’s Elo of 1309. GDPval-AA is our leading metric for general agentic performance, measuring the performance on knowledge work tasks such as preparing presentations and analysis. Models are given code execution and web browsing tools in an agentic loop via our open source reference agentic harness called Stirrup. This continues Kimi K2.6’s strength in tool use, maintaining a 96% score on τ²-Bench Telecom, placing it among other frontier models in this category. ➤ Low hallucination rate: Kimi K2.5 scores 6 on the AA-Omniscience Index, our knowledge evaluation measuring both accuracy and hallucination rate. This score is primarily driven by a comparatively low hallucination rate of 39% (reduced from Kimi K2.5’s 65%), indicating a greater capability to abstain rather than fabricate knowledge when the model is uncertain. Kimi K2.6’s low hallucination rate places it similarly to other models such as Claude Opus 4.7 (36%) and MiniMax-M2.7 (34%) ➤ High token usage: Kimi K2.6 demonstrates high token usage, but is in line with other frontier models in the same intelligence tier. To run the full Artificial Analysis Intelligence Index, Kimi K2.6 used ~160M reasoning tokens. This is slightly lower than Claude Sonnet 4.6 (~190M reasoning tokens) but much higher than GPT 5.4 (~110M reasoning tokens). ➤ Open weights: Kimi K2.6 is a Mixture-of-Experts (MoE) model with 1T total parameters and 32B active, same as the previous two generations of models Kimi K2 Thinking and Kimi K2.5. Kimi K2.6 again pushes the open weights frontier in intelligence. ➤ Third Party Access: Kimi K2.6 is accessible through Moonshot’s First Party API as well as third party API providers Novita, Baseten, Fireworks, and Parasail ➤ Multimodality: Kimi K2.6 supports Image and Video input and text output natively. The model’s max context length remains 256k.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Anthropic's new Claude Opus 4.7 has debuted at the top of the Artificial Analysis Intelligence Index with a score of 57, a 4-point uplift over Opus 4.6 (Adaptive Reasoning, Max Effort, 53). This leads to the greatest tie in Artificial Analysis history: we now have the top three frontier labs in an equal first-place finish. More details in the article below:

  • ImagineArt 2.0 debuts at #9 on our Text to Image Leaderboards, delivering quality comparable to grok-imagine-image from xAI and Imagen 4 Ultra from Google! ImagineArt 2 is the latest proprietary image model from ImagineArt, a popular AI creative studio app that provides users access to various image and video models in one place. ImagineArt 2.0 is currently available as an option in the ImagineArt Image Studio app, with an API for developers coming soon.

    • No alternative text description for this image
  • Fish Audio S2 Pro is the new leading Open Weights model on the Artificial Analysis Speech Arena Leaderboard, closing the gap between Open Weights and Proprietary models Fish Audio S2 Pro is the latest TTS model from Fish Audio, featuring multi-speaker, multi-turn generation and inline prosody and emotion control using natural language tags, e.g., [whisper], [laughing], [excited tone], and [professional broadcast tone]. It was trained on 10M+ hours of audio across 80+ languages, with weights and fine-tuning code available on Hugging Face and an API available on their platform. Key takeaways: ➤ Quality: Fish Audio S2 Pro has an Elo of 1,165, placing it 6th overall and 1st for Open Weights models ahead of other top Open Weights models, Step Audio EditX (1,105), and Kokoro 82M v1.0 (1,057). ➤ Pricing: The model’s API is priced at $15/1M characters via the Fish Audio platform, and is also available for self-hosting via weights on Hugging Face. ➤ Speed: Fish Audio S2 Pro processes 51 characters per second. For comparison, Kokoro 82M v1.0 hosted on Replicate processes 290 characters per second. See more details below ⬇️

    • No alternative text description for this image
  • Artificial Analysis reposted this

    Artificial Analysis released a Model Recommender that ranks models based on your priorities across intelligence, speed, and cost. When you set intelligence as very important, speed as critical, and cost as important, Mercury 2 ranks first. If you're building agents, voice apps, search, or coding tools where latency compounds across every call, the recommender is a useful way to pressure-test your model choice.

    • No alternative text description for this image
  • Artificial Analysis reposted this

    NVIDIA Magpie-TTS just climbed from #8 to #4 on the Artificial Analysis TTS Arena for open-source models. Same 357M parameters. The difference? Switching from parallel codebook prediction to autoregressive Local Transformers. A single architectural change that eliminated audio artifacts and noticeably improved naturalness. 📽️ The attached video is a fun demo where the old and new checkpoints interview each other about what changed. The entire thing, script, voices, images, and video stitching, was AI generated. Both voices produced with Magpie-TTS on a single NVIDIA T4 GPU launched from Brev. 🤗 Model: https://lnkd.in/gaNwUKYc 🗞️ Paper: https://lnkd.in/gem8J5Nv 📊 Leaderboard: https://lnkd.in/gyizEeuQ 🌿 Brev: https://lnkd.in/g_NUF556 #TTS #SpeechAI #OpenSource #VoiceAI

  • Anthropic launched Claude Opus 4.7 today, the new number 1 in our GDPval-AA benchmark for performance on agentic real-world work tasks Opus 4.7 scored 1753 on GDPval-AA at launch with its ‘max’ effort setting, surpassing GPT-5.4 xhigh. This is a significant upgrade, placing Opus back on top of Sonnet on the GDPval-AA leaderboard. Compared to OpenAI’s GPT-5.4, it has an implied win rate of ~60% when compared head-to-head on the GDPval task set. We supported Anthropic with testing this model ahead of release and appreciate them referencing our evaluations in their announcement post and system card for both GDPval-AA and AA-Omniscience. We’re actively conducting the rest of the Artificial Analysis Intelligence Index evaluations and will share complete results soon!

    • No alternative text description for this image
  • Google’s new Gemini 3.1 Flash TTS ranks #2 on the Artificial Analysis Speech Arena Leaderboard, ahead of ElevenLabs’ Eleven v3 and only behind Inworld TTS 1.5 Max Gemini 3.1 Flash TTS represents a significant step forward for Google from previous TTS models, with notably increased naturalness of speech samples. The model now ranks just 4 Elo points behind the leading model on the Speech Arena, the tightest margin at the top of the leaderboard. Key takeaways: ➤ Quality: Gemini 3.1 Flash TTS has an Elo of 1,211 based on over 1.7k arena appearances, placing it just 4 points behind the leading model (Inworld TTS 1.5 Max at 1,215) and 32 points ahead of Eleven v3 at 1,179 ➤ Pricing: Standard pricing of $36.6/1M characters, 3.7x more expensive than Inworld TTS 1.5 Max ($10/1M chars) but 4.7x cheaper than Eleven v3 ($172/1M chars). Expect to be lower for Batch pricing. ➤ Speed: 27.4 characters per second, compared to 138 chars/s for Inworld TTS 1.5 Max and 38.8 chars/s for Eleven v3 ➤ Prompting: Features the ability to generate voices based on text prompting. Google's prompting strategy guide includes elements such as character persona, scene, style, pacing, and accent See more details below.

    • No alternative text description for this image
  • MiniMax M2.7 is now open weights, just over three weeks after launching with a score of 50 in the Artificial Analysis Intelligence Index. However, MiniMax is releasing the model with a non-commercial license. At 230B total with 10B active parameters, M2.7 is ~3.3x smaller than the leading open weights model GLM-5.1 (754B / 40B active) and has ~4x fewer active parameters. This makes M2.7 a highly compelling option for self-deployment, and leads to hosted inference being ~4x cheaper across providers. MiniMax’s decision to restrict commercial usage the M2.7 weights may be part of an emerging trend in how Chinese labs are approaching open source. Proprietary model releases in recent weeks from Chinese labs have included Xiaomi’s MiMo V2 Pro and Alibaba’s Qwen3.6 Plus. Look out for an update on M2.7 providers soon!

    • No alternative text description for this image
    • No alternative text description for this image
  • Revealing HappyHorse-1.0 as the latest video model from Alibaba! HappyHorse-1.0 has landed in #1 or #2 across all of the leaderboards in the Artificial Analysis Video Arena. In our ‘without audio’ leaderboards, HappyHorse-1.0 is comfortably in first place. In our ‘with audio’ leaderboards, its Elo score is almost identical to ByteDance’s Dreamina Seedance 2.0. HappyHorse-1.0 was created by Alibaba-ATH, and it supports all four video generation modalities: Text to Video and Image to Video, each with and without native audio. API access is planned for launch on April 30.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • We’re unveiling a new look for Artificial Analysis! We’ve come a long way since launching Artificial Analysis over 2 years ago. Today, we benchmark 400+ models, 50+ inference providers, and benchmark not only language models but also image, video, speech, music, hardware, and agents. Our mission to support the AI ecosystem with independent benchmarking remains the same, but our brand and website refresh is designed to better reflect how much we’ve grown and how much further we plan to go. A huge thank you to everyone who has been part of the Artificial Analysis community along the way: from developers choosing models and building agents, to labs, inference and hardware providers, and fellow independent researchers. Check out our new look at https://lnkd.in/gVh662hY

  • GLM-5.1 takes the open weights lead on the Artificial Analysis Intelligence Index with a modest gain over GLM-5, with most of the improvement driven by gains on agentic real-world use cases (GDPval-AA) GLM-5.1 is now the leading open weights model in GDPval-AA, ahead of MiniMax-M2.7, and behind GPT-5.4 (xhigh), Claude Opus 4.6 (max) and Claude Sonnet 4.6 (max). Z.ai has now released GLM-5.1’s weights. The model has been available for a few days, but only to subscribers of Zai's Coding Plan. There is no architecture change from GLM-5: GLM-5.1 retains the 744B total / 40B active parameter Mixture-of-Experts design with DeepSeek Sparse Attention, a 200K context window, and BF16 native precision. Since GLM-5, Zai has also released two proprietary models: GLM-5-Turbo, a text-only model that Zai describes as "deeply optimized for the OpenClaw scenario", scoring 47 on the Intelligence Index, and GLM-5V-Turbo (Reasoning), a natively multimodal variant scoring 43 on the Intelligence Index. Both sit below the open weights GLM-5 (Reasoning, 50) and GLM-5.1 (Reasoning, 51) on the Intelligence Index. Key takeaways from benchmarking GLM-5.1 (Reasoning): ➤ GLM-5.1 (Reasoning) scores 51 on the Intelligence Index, a 1 point gain over GLM-5 (Reasoning, 50), and takes the leading open weights position. GLM-5.1 sits ahead of all other open weights models, including MiniMax-M2.7 (50) and Kimi K2.5 (Reasoning, 47), and behind frontier proprietary models including Gemini 3.1 Pro Preview (57), GPT-5.4 (xhigh, 57), and Claude Opus 4.6 (Adaptive Reasoning, max effort, 53) ➤ GDPval-AA is the standout result, with GLM-5.1 reaching an Elo of 1535. This is a +128 Elo gain over GLM-5 (1407) and places GLM-5.1 #4 overall on GDPval-AA, behind only GPT-5.4 (xhigh), Claude Sonnet 4.6 (Adaptive Reasoning, max effort), and Claude Opus 4.6 (Adaptive Reasoning, max effort) ➤ Underlying eval movement is broadly positive, with gains in graduate-level reasoning (GPQA Diamond), instruction following (IFBench), and research-level physics (CritPt). Versus GLM-5 (Reasoning), we observed gains in GPQA Diamond (+4.8 points), IFBench (+4.0 points), CritPt (+2.6 points), and HLE (+0.8 points), with a small regression in SciCode (-2.4 points). TerminalBench Hard, τ²-Bench Telecom, AA-LCR, and AA-Omniscience remain equivalent to GLM-5 ➤ GLM-5.1 is slightly less token efficient than GLM-5, using ~120M output tokens to run the Intelligence Index versus ~109M for GLM-5. Among the open weights peers at the top of the Intelligence Index, GLM-5.1 uses more output tokens than both MiniMax-M2.7 (87M) and Kimi K2.5 (Reasoning, 89M) Key model details: ➤ Context window: 200K tokens, equivalent to GLM-5 ➤ Multimodality: Text input and output only ➤ License: MIT ➤ Availability: GLM-5.1 is available via Zai's first-party API and several third-party providers including Deep Infra Inc., FriendliAI, Novita AI, GMI Cloud, Parasail, Fireworks AI and SiliconFlow

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Artificial Analysis reposted this

    View profile for Bertie Vidgen

    AI Researcher @ Mercor

    Excited to see APEX-Agents on Artificial Analysis 🌟 Their team implemented the benchmark from scratch and made a few small changes to the setup: (1) they run evals with their harness, Stirrup; (2) they dropped two of the IB worlds that use additional services; and (3) they score models based on 3 runs for each task. Interestingly, the scores and rank order of models only changes a little bit from our leaderboard. It's great to see the robustness of the eval -- and where these small differences lead to differences. And if you want to implement APEX-Agents for yourself, you can! It's available open-source on HuggingFace and the inference service is on Github.

    View organization page for Mercor

    696,573 followers

    We are excited to announce our collaboration with Artificial Analysis on APEX-Agents-AA — an independent, live leaderboard evaluating AI agents on the professional tasks that knowledge workers do every day. The leaderboard is built on APEX-Agents, Mercor's open-source benchmark of 480 tasks across investment banking, management consulting, and corporate law — including tool implementations, rubrics, and grading workflows, all available to the community for evaluation and training. Artificial Analysis runs a subset of these tasks through their open-source Stirrup harness, providing a reproducible, independent baseline that any team can verify and build on. APEX-Agents-AA results: 🥇 GPT-5.4: 33.3% 🥈 Claude Opus 4.6: 33.0% 🥉 Gemini 3.1 Pro Preview: 32.0% The top three frontier models are separated by just 1.3 percentage points. The leaderboard will update with key model releases. Check it out at the link in the comments.

    • No alternative text description for this image
  • Artificial Analysis reposted this

    View organization page for Mercor

    696,573 followers

    We are excited to announce our collaboration with Artificial Analysis on APEX-Agents-AA — an independent, live leaderboard evaluating AI agents on the professional tasks that knowledge workers do every day. The leaderboard is built on APEX-Agents, Mercor's open-source benchmark of 480 tasks across investment banking, management consulting, and corporate law — including tool implementations, rubrics, and grading workflows, all available to the community for evaluation and training. Artificial Analysis runs a subset of these tasks through their open-source Stirrup harness, providing a reproducible, independent baseline that any team can verify and build on. APEX-Agents-AA results: 🥇 GPT-5.4: 33.3% 🥈 Claude Opus 4.6: 33.0% 🥉 Gemini 3.1 Pro Preview: 32.0% The top three frontier models are separated by just 1.3 percentage points. The leaderboard will update with key model releases. Check it out at the link in the comments.

    • No alternative text description for this image
  • Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights Muse Spark is a new model from AI at Meta evaluated on Artificial Analysis. We were given early access by Meta to independently benchmark the model. It is the first frontier-class model from Meta since Llama 4 Maverick was released in April 2025, and notably the first Meta model that is not being released as open weights. The release follows Meta's reorganization of its AI efforts under Meta Superintelligence Labs, and signals that Meta is re-entering the frontier race after roughly a year of relative quiet. For context, Llama 4 Maverick and Scout scored 18 and 13 respectively on the Artificial Analysis Intelligence Index as non-reasoning models at the time of their release, while Muse Spark scores 52. Muse Spark essentially closes the gap between to the frontier in a single release. Key takeaways from our benchmarks: ➤ Muse Spark scores 52 on the Artificial Analysis Intelligence Index, placing it within the top 5 models we have benchmarked. It sits ahead of Claude Sonnet 4.6, GLM-5.1, MiniMax-M2.7, Grok 4.20 and behind Gemini 3.1 Pro Preview, GPT-5.4 and Claude Opus 4.6 ➤ Muse Spark is notably token efficient for its intelligence level. It used 58M output tokens to run the Intelligence Index, comparable to Gemini 3.1 Pro Preview (57M) and notably lower than Claude Opus 4.6 (Adaptive Reasoning, max effort, 157M), GPT-5.4 (xhigh, 120M) and GLM-5 (110M) ➤ Muse Spark is the second-most capable vision model we have benchmarked. It scores 80.5% on MMMU-Pro, behind only Gemini 3.1 Pro Preview (82.4%) ➤ Muse Spark performs strongly on reasoning and instruction-following evaluations. It scores 39.9% on HLE, trailing only Gemini 3.1 Pro Preview (44.7%) and GPT-5.4 (xhigh, 41.6%). The model also achieved 5th highest in CritPT with a score of 11%, an eval that is focused on difficult physics research questions. This is substantially above above Gemini 3 Flash (9%) and Claude 4.6 Sonnet (3%) ➤ Agentic performance does not stand out. On GDPval-AA, our evalaution focused on real world work tasks, Muse Spark scores 1427, behind both Claude Sonnet 4.6 at 1648 and GPT-5.4 at 1676, but ahead of Gemini 3.1 Pro Preview at 1320. On On TerminalBench Hard, Muse Spark trails Claude Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro. Muse Spark joins others in achieving a high τ²-Bench Telecom score of 92% Key model details: ➤ License: Proprietary, Meta's first frontier model not released as open weights ➤ Availability: No public API at the time of publishing. Meta expects to provide API access soon. Meta has started integration into their first party AI offering Meta AI and inside Facebook, Instagram, and Threads

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Announcing APEX-Agents-AA, our latest leaderboard on Artificial Analysis, evaluating AI agents on long-horizon professional services tasks with realistic application dependencies This is our implementation of the APEX-Agents benchmark - an agentic work task evaluation open-sourced by Mercor. It tests AI agent ability to execute realistic tasks created by investment banking analysts, management consultants, and corporate lawyers. Mercor released extensive data to enable model evaluation and training across the community, comprising 480 tasks including tool implementations, rubrics, and grading workflows. We exclude tasks with external service dependencies and run the remaining 452 tasks for APEX-Agents-AA. Models complete tasks using Stirrup, our open-source agent harness as used in GDPval-AA, and a customized tool set based on the original benchmark implementation. Results overview: 🥇 OpenAI, Anthropic and Google are in close competition at the top of the leaderboard, with 33.3% for GPT-5.4, 33.0% for Claude Opus 4.6, and 32% for Gemini 3.1 Pro Preview 📈 The overall scores on Artificial Analysis today are similar to Mercor’s testing, but some models such as GPT-5.4 nano show improvements in score using our Stirrup test harness ↻ We’ll be updating this leaderboard with key releases for agentic work use as a metric for agent capability on well-defined, long horizon work tasks APEX-Agents overview: ➤ Tasks span 3 professional domains: investment banking, management consulting, and corporate law ➤ The tasks are designed to require long-horizon work with a large number of tools, which are provided through MCP servers as would be used in many real-world deployments (including calendar, chat, spreadsheet and presentation operations, etc.) ➤ Required outputs include direct message responses (87%) and creating or modifying spreadsheets (6.6%), documents (4.8%), and presentations (1.3%) ➤ Model outputs are parsed and graded against binary rubrics using an LLM judge. Each task is run 3 times and scored pass@1 - a pass requires every rubric test to pass ➤ In our APEX-Agents-AA implementation, 452 tasks run in our open-source Stirrup harness with tool management and usage from Mercor's original MCP implementation. This provides a consistent, reproducible baseline for comparing raw model capability that aligns with realistic agent deployments

    • No alternative text description for this image
  • 🇰🇷 South Korean AI lab Upstage has launched Solar Pro 3! Solar Pro 3 scores 26 on the Artificial Analysis Intelligence Index, a significant improvement over Solar Pro 2 and is currently the second strongest model released by a Korean lab Key benchmarking takeaways: ➤ Strength in agentic tool use and instruction following: Solar Pro 3 scores 71% on IFBench, which signals strong instruction following capabilities. Solar Pro 3 ranks near the frontier models in this category, scoring similarly to GLM-5 (71%) and Kimi K2.5 (70%) and is the leader among Korean models. Solar Pro 3 scores also 86% on τ²-Bench Telecom, demonstrating strong performance on agentic tool-use, making it a strong candidate for incorporation into agentic workflows. ➤ Relatively high token usage: Solar Pro 3 demonstrates relatively high token usage compared to other models in the same intelligence tier, using ~100M reasoning tokens across the Artificial Analysis Intelligence suite. This is comparable to LG’s K-EXAONE (100M reasoning tokens), another Korean model. ➤ Modest accuracy and reliability: Solar Pro 3 scores -54 on AA-Omniscience, our evaluation of knowledge reliability and hallucination, where scores range from -100 to 100 (higher is better) and a negative score indicates more incorrect than correct answers. However, with an 18% on accuracy component score, Solar Pro 3 does outperform Korean competitors in this metric. ➤ First-party API Access: Solar Pro 3 is a proprietary model and is currently available through Upstage’s first-party API Other Relevant Model Details: ➤ Model type: Mixture of Experts (MoE) ➤ Size: 102B total parameters (12B active parameters) ➤ Context length: 128k ➤ Training data cut-off: July 2025

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • We’ve launched agent landscape overviews across 7 key categories relevant to real world tasks agents are used for today! 💼 Categories so far include: General Work, Coding, Chatbots, Presentations, OCR, Data Analysis, and Customer Support. We report on key capabilities relevant to each agent category such as filetype handling, integrations, browser automation, bring-your-own-model support, open source status, and more. This is just a start of our benchmarking of agents. We’ll continue to dive deeper over time with more quantitative analyses.

    • No alternative text description for this image
  • Google has released Gemma 4, four open weights models with multimodality support. The flagship 31B model (39 on the Intelligence Index) uses ~2.5x fewer output tokens than Qwen3.5 27B (Reasoning, 42) but trails it by 3 points on intelligence Google DeepMind Gemma 4 includes four sizes: Gemma 4 31B (dense, 39 on the Intelligence Index), Gemma 4 26B A4B (MoE, 4B active, 31), Gemma 4 E4B (8B, 19), and Gemma 4 E2B (5.1B total, 2.3B active, 15). Gemma 3 was instruct-only at 27B, 12B, 4B, 1B, and 270M; Gemma 4 adds reasoning mode, native video and image support across all sizes (with audio input for Gemma 4 E2B and E4B), doubled context windows, and Apache 2.0 licensing. The nearest open weights models by intelligence to the 31B are Qwen3.5 27B (Reasoning, 42), GLM-4.7 (Reasoning, 42), MiniMax-M2.5 (42), and DeepSeek V3.2 (Reasoning, 42). Qwen3.5 also supports images and video natively; DeepSeek V3.2 and MiniMax-M2.5 are text-only. Key benchmarking results for the reasoning variants: ➤Gemma 4 represents a large intelligence jump over Gemma 3. Gemma 4 31B (Reasoning, 39) is +29 points over Gemma 3 27B Instruct (10), Gemma 4 E4B (19) is +13 points over Gemma 3n E4B Instruct (6), and Gemma 4 E2B (15) is +10 points over Gemma 3n E2B Instruct (5). Context windows also doubled from 128K to 256K for the larger models, and increased 4x from 32K to 128K for E2B and E4B ➤ Gemma 4 31B (Reasoning, 39) trails Qwen3.5 27B (Reasoning, 42) by 3 points, primarily due to weaker agentic performance. On non-agentic evaluations, the models are more competitive: Gemma 4 31B leads on SciCode (43% vs 40%) and TerminalBench Hard (36% vs 33%), while scoring similarly on GPQA Diamond (86% vs 86%), IFBench (76% vs 76%), and HLE (23% vs 22%) ➤ Gemma 4 31B is notably token efficient, using 39M output tokens to run the Intelligence Index vs 98M for Qwen3.5 27B (Reasoning). This is ~2.5x fewer output tokens for a model scoring 3 points lower. For context, the other models at the 42-point intelligence level also use significantly more tokens: MiniMax-M2.5 (56M), DeepSeek V3.2 (Reasoning, 61M), and GLM-4.7 (Reasoning, 167M) ➤ Gemma 4 26B A4B (Reasoning, 31) activates just 4B of its 27B total parameters and is ahead of select peers in the ~3-4B active parameter range. Qwen3.5 35B A3B (Reasoning, 37) leads models with ~3B active parameters and is 6 points ahead of Gemma 4 26B A4B, with notably stronger agentic capabilities (Agentic Index 44 vs 32). GLM-4.7-Flash (Reasoning, 30) scores slightly lower than Gemma 4 26B A4B with 3B active parameters Key model details: ➤ Context window: 256K tokens (31B, 26B A4B), 128K tokens (E4B, E2B). ➤Multimodality: All models support text, images, and video input. E2B and E4B also support native audio input ➤ API availability: The two larger models are available for free on Google AI Studio. There are several third-party providers hosting the larger Gemma 4 variants such as Novita, LightningAI, and Parasail

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Dreamina Seedance 2.0 from ByteDance Seed takes the #1 spot across all modalities in the Artificial Analysis Video Arena, surpassing Kling 3.0, Grok Imagine, and Veo 3.1! Dreamina Seedance 2.0 is the latest video generation model from ByteDance Seed, capable of generating videos up to 15 seconds with native stereo audio support. It also accepts text, images, and video as inputs, including multiple image references in a single generation. Dreamina Seedance 2.0 is currently available for whitelisted customers on the Dreamina AI app, with general availability coming later. See example generations of Dreamina Seedance 2.0 in the Artificial Analysis Video Arena 🧵

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • India enters the open-weights AI race with its largest models pre-trained from scratch: Sarvam 105B and Sarvam 30B Sarvam's Sarvam 105B and Sarvam 30B score 18 and 12 on the Artificial Analysis Intelligence Index respectively. Announced at the India AI Impact Summit 2026 and open-sourced under Apache 2.0, both are Mixture-of-Experts models trained entirely in India using compute provided under the IndiaAI Mission. Both support reasoning and non-reasoning modes. These are an improvement from Sarvam's previous model, Sarvam M (8 on Intelligence Index, 23.6B parameters), which was based on Mistral Small rather than pre-trained from scratch. Sarvam 105B has 106B total parameters with ~10B active per token and a 128K context window. Sarvam 30B has 32B total parameters with ~2.4B active per token and a 65K context window. Alongside the text models, Sarvam also announced Saaras v3 (Speech to Text) and Bulbul v3 (Text to Speech) with a focus on Indic languages. Key takeaways in reasoning mode: ➤ Sarvam 105B scores 18 on the Intelligence Index. Among ~100B-class open-weights reasoning models, it trails GLM-4.5-Air (23), INTELLECT-3 (22), Mistral Small 4 (27), and gpt-oss-120B (High, 33). All four peers also activate more parameters per token. ➤ Sarvam 30B scores 12 on the Intelligence Index. Among ~30B-class open-weights reasoning models, it trails GLM-4.7-Flash (30), Nemotron Cascade 2 30B A3B (28), Qwen3 30B A3B 2507 (22), and Qwen3 32B (17). Sarvam 30B activates fewer parameters than these peers. ➤ Sarvam 105B's relative strength is in select agentic tasks. Its agentic index of 25 places it ahead of INTELLECT-3 (20) and GLM-4.5-Air (21) despite trailing both on overall intelligence. Its GDPval index of 773 also edges ahead of GLM-4.5-Air (665). Both new models are a large step up from Sarvam M (Reasoning), which scored 8 on the Intelligence Index. ➤ Compared to peers, both models score lower on TerminalBench Hard (Agentic Coding & Terminal Use) and AA-Omniscience. Sarvam 105B scored 1.5% and Sarvam 30B scored 2.3% on TerminalBench Hard, compared to GLM-4.5-Air (20.5%) and INTELLECT-3 (9.1%). The AA-Omniscience Index is -60 for Sarvam 105B and -72 for Sarvam 30B. Both models have high hallucination rates relative to their accuracy, and both attempt to answer far more questions rather than abstaining, which drives the negative scores. Key model details: ➤ Modality: Text input and output only. ➤ Context window: 128K tokens (Sarvam 105B) and 65K tokens (Sarvam 30B). ➤ Pricing: Currently free on Sarvam's first-party API. ➤ License: Apache 2.0. ➤ Availability: Sarvam's first-party API; weights available on Hugging Face and AIKosh. See how the models compare to other models you are using: https://lnkd.in/gVh662hY

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Microsoft has released MAI-Transcribe-1: a speech transcription model achieving 3.0% on AA-WER (#4), and is fast at 69x realtime The model was developed by Microsoft AI (MAI)’s Superintelligence team and supports 25 languages including English, French, Arabic, Japanese, and Chinese. MAI-Transcribe-1 API is currently available in public preview via Azure Speech on Microsoft Foundry. On the Artificial Analysis Speech to Text (STT) leaderboard, MAI-Transcribe-1 achieves a 3.0% word error rate on AA-WER for speech transcription accuracy, positioning it 4th overall behind Mistral’s Voxtral Small (2.9% AA-WER), Google’s Gemini 3.1 Pro High (2.9% AA-WER) and ElevenLabs’ Scribe v2 (2.3% AA-WER). It also stands out as one of the faster high-accuracy transcription models available, processing audio at ~69x real-time. See more details below ⬇️

    • No alternative text description for this image
  • Google has released Gemma 4, a new family of multimodal open-weight models including Gemma 4 E2B, Gemma 4 E4B, Gemma 4 31B and Gemma 4 26B A4B Google DeepMind's new Gemma 4 family introduces four multimodal models supporting text, image, and video inputs. We evaluated Gemma 4 31B (dense) and Gemma 4 26B A4B (MoE), both with a 256k context window, while the other two smaller models support up to 128k. With 31B and 26B parameters respectively, both evaluated models can run on a single H100. On GPQA Diamond, our scientific reasoning evaluation, Gemma 4 31B (Reasoning) scores 85.7%, the second highest result we have recorded for an open-weights model with fewer than 40B parameters, just behind Qwen3.5 27B (Reasoning, 85.8%). It reaches this score using only ~1.2M output tokens, fewer than Qwen3.5 27B (~1.5M) and Qwen3.5 35B A3B (~1.6M). Gemma 4 26B A4B (Reasoning) scores 79.2%, ahead of gpt-oss-120B (high, 76.2%) but behind Qwen3.5 9B (Reasoning, 80.6%). We are now running the Artificial Analysis Intelligence Index on all four Gemma 4 models and will share a full update once those results are complete.

    • No alternative text description for this image
    • No alternative text description for this image
  • KwaiKAT has released KAT-Coder-Pro V2, a non-reasoning model that scores 44 on the Artificial Analysis Intelligence Index, an 8 point improvement from KAT-Coder-Pro V1 KwaiKAT has updated their flagship proprietary coding model with the release of KAT-Coder-Pro V2. KAT-Coder-Pro V2 achieves 44 on the Artificial Analysis Intelligence Index, matching Claude Sonnet 4.6 (non-reasoning) and trailing only Claude Opus 4.6 (non-reasoning, 46) among non-reasoning models. At ~9M output tokens, it is also more token efficient than Claude Opus 4.6 (~11M), Claude Sonnet 4.6 (~14M), and reasoning models with similar intelligence such as DeepSeek V3.2 (reasoning, ~61M) and Qwen3.5 397B A17B (reasoning, ~86M). KAT-Coder-Pro V2 is a non-reasoning model, unlike all of the current frontier language models which ‘think’ before answering. Typically, reasoning variants score higher on the Intelligence Index than their non-reasoning counterparts, but consume more output tokens and are less suited to latency-sensitive workloads. More details in the article below.

  • Cohere has released Cohere Transcribe: an open weights model achieving 4.7% on AA-WER, based on 3 datasets including our proprietary AA-AgentTalk dataset The 2B parameter model is based on a conformer encoder-decoder architecture. It was trained from scratch on 14 languages including English, French, Mandarin, Japanese, and Arabic. On the Artificial Analysis Speech to Text (STT) leaderboard, Cohere Transcribe achieves a 4.7% word error rate on AA-WER for speech transcription accuracy, positioning it near NVIDIA's Canary Qwen 2.5B (4.4% AA-WER, 2.5B parameters) and OpenAI's Whisper Large v3 (4.2% AA-WER, 1.6B parameters). It is also among the faster transcription models available, processing ~60 seconds of audio in approximately one second. Cohere Transcribe is currently available for free (subject to rate limits) via Cohere’s API. The model is also available for download on Hugging Face under Apache 2.0 license See more details below ⬇️

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Introducing AA-AgentPerf - the hardware benchmark for the agent era. Key details: ➤ Real agent workloads, not synthetic queries: we’ve captured real coding agent trajectories where our agents used up to 200 turns and worked with sequence lengths >100K tokens ➤ Production optimizations allowed: KV cache reuse, disaggregated prefill/decode, speculative decoding - we’re allowing the optimizations that labs and inference providers are serving in production so that we can capture what real deployments should look like ➤ Measures what developers need to know: Max concurrent users at each target output speed, expressed per accelerator, per kW TDP, per $/hr, and per rack ➤ Built for every kind of scale: designed to measure systems from a single accelerator up to a full rack, and to fairly evaluate every architecture from DRAM-only designs to SRAM-only designs and everything in between ➤ Live now: we’re announcing AA-AgentPerf today and opening submissions of configurations for benchmarking effective immediately. The models supported at launch are gpt-oss-120b and DeepSeek V3.2. We’ll be publishing results on a rolling basis. AA-AgentPerf is a benchmark for real-world performance of AI accelerator hardware. We’re benchmarking inference of particular models on a specific system with a specific config (ie. inference stack, parallelism config and more). AA-AgentPerf has been shaped by our work with inference providers and engagement with AI accelerator companies, developers, and enterprise buyers over the past year. Our goal is for anyone deploying models - whether buying or leasing accelerators - to be able to use AA-AgentPerf as the definitive resource for understanding real-world hardware performance. AA-AgentPerf results will primarily be expressed as a maximum number of concurrent users serviceable at a given per-user token output speed (and vice versa). We will combine these results with several dimensions that are important to developers and customers: ➤ Users per accelerator: the most basic view - how many users can be serviced by a single accelerator at each output speed. ➤ Users per kW TDP: provides context on how power-efficient the accelerators are. ➤ Users per unit rental cost: provides context on how cost-efficient each accelerator is. ➤ Users per rack: provides context on how space-efficient the accelerators are. We expect initial results to be available within the next 1-2 weeks, after submissions from hardware providers and QA from our team. Results will be visible at https://lnkd.in/gR4hmEgY Get to know the evaluation methodology more closely at https://lnkd.in/gsJZYCpp

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Google has released Gemini 3.1 Flash Live Preview, achieving #2 in our Big Bench Audio Speech to Speech model benchmark, and now features configurable thinking levels With thinking level set to high, it scores 95.9% on Big Bench Audio, making it the second-highest scoring speech reasoning model behind Step-Audio R1.1 Realtime (97.0%) and ahead of Grok Voice Agent (92.9%). Switching to minimal thinking brings the score down to 70.5%, but opens up a faster option for latency-sensitive applications. The flexibility in thinking levels also provides a range of latency profiles. On high, average Time to First Audio (TTFA) is 2.98 seconds, slower than Step-Audio R1.1 Realtime (1.51s) and Grok Voice Agent (0.78s). On minimal, TTFA drops to 0.96 seconds, closer to the pack but still behind Google's own Gemini 2.5 Flash Native Audio Dialog (0.63s), which trades ~5 points of intelligence for the fastest response time on our leaderboard. Key takeaways: ➤ Model introduces configurable thinking levels (minimal, low, medium, high) that let developers dial reasoning depth up or down ➤ "High" thinking level: 95.9% Big Bench Audio score (2nd overall, behind only Step-Audio R1.1 Realtime ), 2.98s TTFA ➤ "Minimal" thinking level: 70.5% score, 0.96s TTFA ➤ Pricing remains stable at $0.35 per hour of audio input, and $1.38 per hour audio output, matching Gemini 2.5 Flash Native Audio Dialog See below for more details 👇

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Inworld, ElevenLabs, and MiniMax continue to lead our Text to Speech leaderboard for most preferred models Recent checkpoints from each of the labs continue to push the frontier of TTS quality, with 4 out of the top 5 models being released this year. Leading TTS models are increasingly realistic, particularly on relatively straightforward text, with preference differences increasingly coming down to affinity for different voices. Latest results also reflect stronger bot vote filtering, confirmed via triangulation against third-party evaluators. We've also added rank ranges based on each model's 95% confidence interval, showing where a model could land based on its Elo score range. Key results: ➤ Most preferred: Current top 5 per our TTS leaderboard: 1. Inworld TTS 1.5 Max (Elo of 1,238); 2. ElevenLabs Eleven v3 (1,197); 3. Inworld TTS 1 Max (1,183); 4. Inworld TTS 1.5 Mini (1,182); 5. MiniMax Speech 2.8 HD (1,175) ➤ Price: Kokoro 82M v1.0 (Replicate) leads at $0.65 per 1M characters, followed by Inworld TTS 1 and 1.5 Mini at $5, and AsyncFlow V2 at $8.33 ➤ Speed: WaveNet leads for batch generation at 419 characters processed per second, followed by Kokoro 82M v1.0 (Replicate) at 235, and Inworld TTS 1.5 Mini at 214 See below for further detail ⬇️

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Artificial Analysis reposted this

    View profile for George Cameron

    Co-Founder at Artificial Analysis

    Had the pleasure of a 20 minute meeting with Jensen at NVIDIA GTC last week! We discussed the importance of independent benchmarks and how Artificial Analysis can further support the AI ecosystem, from agents to models to hardware. Our benchmarks were also referred to three times in Jensen’s keynote. He also gave Artificial Analysis a NVIDIA Blackwell GB200 chip signed 'The benchmark of excellence'. The only problem is that we might blow a fuse if we try plugging it in at our office. 😆 Certainly a GTC to remember!

    • No alternative text description for this image
  • Mistral has released Mistral Small 4, an open weights model with hybrid reasoning and image input, scoring 27 on the Artificial Analysis Intelligence Index Mistral AI's Small 4 is a 119B mixture-of-experts model with 6.5B active parameters per token, supporting both reasoning and non-reasoning modes. In reasoning mode, Mistral Small 4 scores 27 on the Artificial Analysis Intelligence Index, a 12-point improvement from Small 3.2 (15) and now among the most intelligent models Mistral has released, surpassing Mistral Large 3 (23) and matching the proprietary Magistral Medium 1.2 (27). However, it lags open weights peers with similar total parameter counts such as gpt-oss-120B (high, 33), NVIDIA Nemotron 3 Super 120B A12B (Reasoning, 36), and Qwen3.5 122B A10B (Reasoning, 42). Read more in the article below.

  • MiniMax has released MiniMax-M2.7, delivering GLM-5-level intelligence for less than one third of the cost MiniMax-M2.7 from @MiniMax_AI scores 50 on the Artificial Analysis Intelligence Index, an 8-point improvement over MiniMax-M2.5, which was released one month ago. This is driven by stronger performance on real-world agentic tasks and reduced hallucinations. MiniMax-M2.7 is now ahead of MiMo-V2-Pro (Reasoning, 49) and Kimi K2.5 (Reasoning, 47), and equivalent to GLM-5 (Reasoning, 50) while using 20% fewer output tokens and costing less than a third as much to run. MiniMax-M2.7 is a reasoning-only model and maintains the same per-token pricing as MiniMax-M2.5. Key takeaways: ➤ Strong performance on real-world agentic tasks: MiniMax-M2.7 achieves a GDPval-AA Elo of 1494, a significant improvement from MiniMax-M2.5 (1203) and ahead of MiMo-V2-Pro (Reasoning, 1426), GLM-5 (Reasoning, 1406), and Kimi K2.5 (Reasoning, 1283). It remains behind frontier models such as GPT-5.4 (xhigh, 1667) and Claude Opus 4.6 (Adaptive Reasoning, max effort, 1606) ➤ Reduced hallucinations: MiniMax-M2.7 scores +1 on the AA-Omniscience Index, up from MiniMax-M2.5 (-40). This is competitive with GPT-5.2 (xhigh, -1) and GLM-5 (Reasoning, +2), and well ahead of Kimi K2.5 (Reasoning, -8). The improvement from M2.5 is purely driven by reduced hallucinations, meaning the model is more likely to abstain from answering when it doesn’t know the answer, rather than guessing. M2.7 achieves a hallucination rate of 34%, lower than Claude Sonnet 4.6 (Adaptive Reasoning, max effort, 46%) and Gemini 3.1 Pro Preview (50%). ➤ Gains across most evaluations compared to MiniMax-M2.5: Outside of the GDPval-AA and AA-Omniscience improvements noted above, MiniMax-M2.7 improves in HLE (+9 p.p.), TerminalBench Hard (+5 p.p.), SciCode (+4 p.p.), IFBench (+4 p.p.), GPQA (+3 p.p.), and LCR (+3 p.p.). We saw a notable regression in τ²-Bench (-11 p.p.). ➤ Increased token use: MiniMax-M2.7 used ~87M output tokens to run the Artificial Analysis Intelligence Index, up 55% from MiniMax-M2.5 (~56M). It remains more token-efficient than other models such as GLM-5 (Reasoning, 110M) and Kimi K2.5 (Reasoning, ~89M) ➤ Leading cost efficiency: MiniMax-M2.7 cost $176 to run the Artificial Analysis Intelligence Index, maintaining the same $0.30/$1.20 per 1M input/output pricing as M2.5. This places it on the Pareto frontier of our Intelligence vs. Cost chart. For context, GLM-5 (Reasoning) cost $547 at equivalent intelligence, Kimi K2.5 (Reasoning) cost $371, and Gemini 3 Flash Preview (Reasoning) cost $278 Key model details: ➤ Context window: 200K tokens (equivalent to MiniMax-M2.5). ➤ Pricing: $0.30/$1.20 per 1M input/output tokens (unchanged from MiniMax-M2.5). ➤ Availability: MiniMax first-party API only. ➤ Modality: Text input and output only (no multimodality). ➤ Licensing: MiniMax has not announced whether MiniMax-M2.7 will be open weights. MiniMax-M2.5 is available under the MIT license.

    • No alternative text description for this image

Similar pages

Browse jobs