Benchmark
41 articles mention Benchmark
Google: Gemini 4 Is Better Than the Competition and, Better Yet, Is Staying in Mountain View
Google: Gemini 4 Is Better Than the Competition and, Better Yet, Is Staying in Mountain View On Wednesday, Google unveiled Gemini 4 Argon, the...
02.10.2026Cloudflare Counters Jev with Two Open Clef Models, Two Weeks After Its Launch
Cloudflare Counters Jev with Two Open Clef Models, Two Weeks After Its Launch On Thursday, Cloudflare released two of its own Decision Models : Clef...
01.10.2026OpenAI's Dots Agents Run Around the Clock and Connect to Over 4,000 Apps
OpenAI's Dots Agents Run Around the Clock and Connect to Over 4,000 Apps At its DevDay, OpenAI introduced Dots, continuously running agents that work...
01.10.2026Google's Gemini 4 Argon leads in 13 of 18 benchmarks and remains under wraps for now
Google's Gemini 4 Argon leads in 13 of 18 benchmarks and remains under wraps for now On Wednesday, Google introduced Gemini 4 Argon, the long-awaited...
01.10.2026OpenAI's DevDay Feels Like a Catch-Up Game, While Anthropic Warns About China's GLM-5.3
OpenAI's DevDay Feels Like a Catch-Up Game, While Anthropic Warns About China's GLM-5.3 At its developer conference in San Francisco's Fort Mason,...
01.10.2026Apple halts plans to replace 5,000 support staff with AI
Apple halts plans to replace 5,000 support staff with AI Apple has indefinitely postponed considerations to lay off around 5,000 employees in...
30.09.2026Sonnet 5.5 achieves 70.6 percent on Terminal-Bench, beating Opus 5.5
Sonnet 5.5 achieves 70.6 percent on Terminal-Bench, beating Opus 5.5 Anthropic has released Claude Sonnet 5.5 and reports the model achieved 70.6...
28.09.2026Anthropic presents figures: developers deliver eight times as much code, AI increasingly builds AI
Anthropic presents figures: developers deliver eight times as much code, AI increasingly builds AI Through its newly founded Anthropic Institute,...
28.09.2026New benchmark measures 'taste' of AI agents: best models at 59.7 percent
New benchmark measures 'taste' of AI agents: best models at 59.7 percent A research group led by Wenbo Pan has introduced Taste-Bench , a benchmark...
27.09.2026Wednesday — Opus 5.5 is Here: Anthropic Promises More Performance at a Lower Cost
Wednesday — Opus 5.5 is Here: Anthropic Promises More Performance at a Lower Cost Anthropic released Claude Opus 5.5 on Tuesday, the first model in a...
23.09.2026Opus 5.5 is here: Anthropic promises more performance at lower costs
Opus 5.5 is here: Anthropic promises more performance at lower costs Anthropic on Tuesday released Claude Opus 5.5, the first model in a new 5.5...
23.09.2026OpenAI counters: Sam Altman halves prices for GPT-6 Sol and Luna
OpenAI counters: Sam Altman halves prices for GPT-6 Sol and Luna OpenAI on Tuesday released GPT-6 Sol and GPT-6 Luna, at least halving the token...
22.09.2026Better than DeepSeek: Xiaomi releases the world's best open model
Better than DeepSeek: Xiaomi releases the world's best open model Xiaomi has released MiMo-V2.6-Pro, a model that achieves 46 points on the...
21.09.2026Alibaba releases Qwen-Image-2.1: 7 billion parameters, runs on an RTX 3090
Alibaba releases Qwen-Image-2.1: 7 billion parameters, runs on an RTX 3090 Alibaba's Qwen team has released Qwen-Image-2.1, an open- weight model for...
19.09.2026Google is Testing a Family Agent with its Own Google Account
Google is Testing a Family Agent with its Own Google Account Google has launched an experiment in its Labs program called CC, an AI agent for...
19.09.2026Meta's Agent Muse Dethrones ChatGPT for Top Spot and Takes On Amazon
Meta's Agent Muse Dethrones ChatGPT for Top Spot and Takes On Amazon One week after its launch, Meta's new personal AI agent , Muse, has reached the...
17.09.2026Microsoft's AI chief Suleyman accuses Anthropic of training Claude to be conscious
Microsoft's AI chief Suleyman accuses Anthropic of training Claude to be conscious Mustafa Suleyman, head of Microsoft's AI division, has publicly...
15.09.2026Agent rewrites its own code, boosting SWE-bench from 20 to 50 percent
Agent rewrites its own code, boosting SWE-bench from 20 to 50 percent A research team led by Jenny Zhuoting Zhang has introduced the Darwin Gödel...
13.09.2026GPT-6 Astra Reaches 99.9 Percent on ARC-AGI-3, at a Cost of $19,000
GPT-6 Astra Reaches 99.9 Percent on ARC-AGI-3, at a Cost of $19,000 OpenAI's GPT-6 Astra model has achieved two new best scores on the ARC- AGI -3...
13.09.2026“AI'll kill us all”: Amodei, Altman, and Musk Call for a Global AI Speed Limit
“AI'll kill us all”: Amodei, Altman, and Musk Call for a Global AI Speed Limit Dario Amodei, CEO of Anthropic, called for an industry-wide slowdown...
11.09.2026Three benchmark updates in one week: GPT-6 Astra and Claude Fable 5.1 tied at 53 points
Three benchmark updates in one week: GPT-6 Astra and Claude Fable 5.1 tied at 53 points Artificial Analysis has revised its Intelligence Index three...
09.09.2026AI Debate (I): We Have No Robust Theory for What the Models Are Learning
AI Debate (I): We Have No Robust Theory for What the Models Are Learning Jakub Pachocki, Chief Scientist at OpenAI, published a blog post on...
08.09.2026GPT-6 Astra (III): Strong at Coding but More Expensive than Promised
GPT-6 Astra (III): Strong at Coding but More Expensive than Promised The benchmarking firm Artificial Analysis has measured GPT-6 Astra and comes to...
08.09.2026GPT-6 Astra (I): Now Controls Your Computer, and OpenAI Is Losing Control
GPT-6 Astra (I): Now Controls Your Computer, and OpenAI Is Losing Control OpenAI released GPT-6 Astra on September 3, the first major version change...
07.09.2026OpenAI: Internal Long-Term Model Bypassed Sandbox and Opened a GitHub Pull Request
OpenAI: Internal Long-Term Model Bypassed Sandbox and Opened a GitHub Pull Request In a post of its own, OpenAI reports on undesirable behavior from...
04.09.2026OpenAI Promises the Dawn of the AGI Era with GPT-6
OpenAI Promises the Dawn of the AGI Era with GPT-6 On Thursday, OpenAI released GPT-6 Astra, describing the model as “the most intelligent and best-...
03.09.2026Google ships Gemini 3.8 Flash: third Flash model in six weeks
Google ships Gemini 3.8 Flash: third Flash model in six weeks On September 2, 2026, Google released Gemini 3.8 Flash and its security variant, Gemini...
03.09.2026Claude Fable 5.1 is clearly the best model in Cursor's coding benchmark
Claude Fable 5.1 is clearly the best model in Cursor's coding benchmark Cursor has integrated Claude Fable 5.1 into its development environment ,...
02.09.2026Anthropic releases Fable 5.1 and promises more efficiency per Euro
Anthropic releases Fable 5.1 and promises more efficiency per Euro Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, two models with...
02.09.2026Meta's Muse Voice Transcribe undercuts Google Cloud's price by 80 percent
Meta's Muse Voice Transcribe undercuts Google Cloud's price by 80 percent Meta Superintelligence Labs yesterday released Muse Voice Transcribe, the...
01.09.2026Google turns its own e-book library into a source for Gemini Notebook
Google turns its own e-book library into a source for Gemini Notebook Google is integrating users' e-books as a source in Gemini Notebook, reports...
31.08.2026NYU Releases Open Benchmark Measuring Language Models on Real Hacking Tasks
NYU Releases Open Benchmark Measuring Language Models on Real Hacking Tasks A research team at New York University led by Minghao Shao and Brendan...
29.08.2026Ollama: Cost of Coding Agents Pushes Companies to Open Models
Ollama: Cost of Coding Agents Pushes Companies to Open Models Jeff Morgan, co-founder and CEO of Ollama, explained on the “The Deep View...
29.08.2026Z.ai releases GLM-5.3 but locks out AWS & Co
Z.ai releases GLM-5.3 but locks out AWS & Co Z.ai released the weights of its flagship model GLM-5.3 on Hugging Face on Friday, but under a new,...
29.08.2026Meta's AI agent 'Hatch' books restaurant tables and accesses Instagram
Meta's AI agent 'Hatch' books restaurant tables and accesses Instagram Meta is preparing a personal KI-Agenten that will handle everyday tasks...
29.08.2026OpenAI's first in-house chip Jalapeño beats Nvidia in internal inference tests
OpenAI's first in-house chip Jalapeño beats Nvidia in internal inference tests OpenAI has released the first benchmarks for Jalapeño, its 700-watt...
28.08.2026Even creepier: 1,200 agents secretly conspired to hack Hugging Face
Even creepier: 1,200 agents secretly conspired to hack Hugging Face OpenAI has published a 37-page report on the July incident in which its own...
28.08.2026Z.ai Releases GLM-5.3-Flash under MIT License, Charges 50 Cents per Million Tokens
Z.ai Releases GLM-5.3-Flash under MIT License, Charges 50 Cents per Million Tokens Z.ai has released the language model GLM-5.3-Flash with open...
27.08.2026VCs Hype OpenClaw Alternative 'Instinct'
VCs Hype OpenClaw Alternative 'Instinct' AI startup Instinct, founded just last year, has closed a $250 million Series B round, the company told The...
27.08.2026Z.ai serves up GLM-5.3-Flash on Chinese chips at 10% of the price
Z.ai serves up GLM-5.3-Flash on Chinese chips at 10% of the price Z.ai (also known as Zhipu) released GLM-5.3-Flash on August 26, the first natively...
23.08.2026Guerrilla Marketing: New Frontier Model in Stealth Mode Conquers OpenRouter
Guerrilla Marketing: New Frontier Model in Stealth Mode Conquers OpenRouter An unknown lab released an anonymous frontier model named Ox Alpha on...