KO EN

AX BRIEF

AI news, benchmarks & engineering blog curation

● LIVE
지능1. Claude Fable 5.1 (max with fallback) 100 pt2. Claude Fable 5 (with fallback) 96 pt3. Claude Opus 5 (max) 95 pt4. GPT-6 Astra (max) 95 pt5. GPT-5.6 Sol (max) 91 pt코딩1. Claude Fable 5.1 (max) 100 pt2. Claude Opus 5 (xhigh) 94 pt3. GPT-6 Astra (max) 91 pt4. Muse Spark 1.3 (xhigh) 81 pt5. Grok 4.5 (high) 81 pt이미지1. GPT Image 2 (high) 100 pt2. MAI-Image-2.6 93 pt3. GPT Image 1.5 (high) 91 pt4. Reve 2.1 88 pt5. Muse Image 85 pt비디오1. Gemini Omni Flash 100 pt2. Minimax H3 Max (post-trained by fal) 94 pt3. Dreamina Seedance 2.0 720p 91 pt4. gemini-omni-1.1-flash 85 pt5. Wan 3.0 85 pt가격1. Claude Fable 5.1 (max with fallback) 100 pt2. GPT-6 Astra (max) 100 pt3. Claude Fable 5 (with fallback) 100 pt4. Claude Opus 5 (max) 49 pt5. GPT-5.6 Sol (max) 39 pt속도1. Gemini 3.5 Flash-Lite 100 pt2. Gemini 3.8 Flash (high) 85 pt3. Muse Spark 1.3 (max) 47 pt4. gpt-oss-120b (high) 43 pt5. GPT-5.6 Luna (max) 28 pt

AI News

Daily AI industry news — funding, products, policy, and major moves from global AI companies and startups.

The Agent Loop: How Claude Code Shifts AI from Drafting to Iterating

The Agent Loop: How Claude Code Shifts AI from Drafting to Iterating

Anthropic's Claude Code introduces agent loops for autonomous iteration. Non-deterministic loops replace traditional recursive programming l

AI Agent Learning Systems: Why Organizational Memory Trumps Model Power

AI Agent Learning Systems: Why Organizational Memory Trumps Model Power

AI agents often repeat errors due to a lack of organizational memory. Agentic learning systems use observability and data fabrics to store k

The Claude AI Pipeline That Found $500,000 in Google API Flaws

The Claude AI Pipeline That Found $500,000 in Google API Flaws

A researcher used Claude AI to automate the discovery of Google API vulnerabilities. The system earned $500,000 in bug bounties within three

Oracle and Meta's AI Pivot: The New Formula for Big Tech Layoffs

Oracle and Meta's AI Pivot: The New Formula for Big Tech Layoffs

Major tech firms are cutting thousands of jobs as AI increases operational speed. Companies are redirecting payroll budgets toward AI infras

Why Open LLMs are Triggering a Linux-Style Shift from GPT and Claude

Why Open LLMs are Triggering a Linux-Style Shift from GPT and Claude

Open LLMs are becoming viable alternatives to GPT and Claude. The performance gap has narrowed to a few months of development. Model soverei

Slack's Agentic Testing Experiment: MCP Cuts Failure Rates to 0-12%

Slack's Agentic Testing Experiment: MCP Cuts Failure Rates to 0-12%

Slack tested agentic E2E testing using MCP, CLI, and generated code. The MCP approach achieved failure rates between 0% and 12%. High costs

OpenAI's Patch the Planet Initiative Secures Open Source With AI

OpenAI's Patch the Planet Initiative Secures Open Source With AI

OpenAI and Trail of Bits launched the Patch the Planet initiative. The project uses Codex Security to find and fix open source bugs. This de

Groq Pivots to NeoCloud After $650 Million Funding and Nvidia Talent Drain

Groq Pivots to NeoCloud After $650 Million Funding and Nvidia Talent Drain

Groq secured $650 million in funding to pivot toward a NeoCloud service model. Nvidia integrated Groq's LPU IP into the new Nvidia Groq 3 LP

Why Claude Code's Extended Thinking Is a Summary, Not a Raw Trace

Why Claude Code's Extended Thinking Is a Summary, Not a Raw Trace

Anthropic reveals that Claude Code's extended thinking is a summary of the reasoning process. Raw thinking blocks are stored in local disk l

Nvidia's Closed-Loop Cooling System Targets Zero On-Site Water Use

Nvidia's Closed-Loop Cooling System Targets Zero On-Site Water Use

Nvidia introduced a closed-loop cooling system to eliminate on-site water consumption. The system uses a 45°C to 55°C thermal cycle to remov

Why Alibaba's HappyHorse 1.1 is Filling the Void Left by Sora

Why Alibaba's HappyHorse 1.1 is Filling the Void Left by Sora

Alibaba Cloud launched HappyHorse 1.1 with a focus on character consistency. The model ranks second on the Arena.ai leaderboard with 1,444 p

Amazon Alexa+ Targets 600 Million Hindi Speakers in India

Amazon Alexa+ Targets 600 Million Hindi Speakers in India

Amazon is beta testing Hindi support for Alexa+ in India. The service is free for Prime members but requires a subscription for others. Amaz

Reflection AI Challenges Closed-Source Giants With SpaceX GB300 Deal

Reflection AI Challenges Closed-Source Giants With SpaceX GB300 Deal

Reflection AI partnered with SpaceX to access Nvidia GB300 chips. The company will pay 150 million dollars monthly for compute. Open-weight

The 93.2 LiveCodeBench Score That Defines Sakana AI's Fugu Ultra

The 93.2 LiveCodeBench Score That Defines Sakana AI's Fugu Ultra

Sakana AI released Fugu as a multi-agent orchestration system to prevent vendor lock-in. Fugu Ultra outperforms Claude models on LiveCodeBen

agent-connector Collapsed 71 Hook Scripts Into a Single Definition

agent-connector Collapsed 71 Hook Scripts Into a Single Definition

The agent-connector tool automates MCP server deployment across 42 agent CLIs. It reduces deployment code from 20,322 lines down to only 76

Self-Harness Boosts LLM Agent Performance by Up to 60%

Self-Harness Boosts LLM Agent Performance by Up to 60%

Shanghai AI Lab introduced Self-Harness to automate agent optimization. The framework improved model performance by 33 to 60 percent. It rep

Agent-Blackbox Cuts Claude Code Token Waste by 44%

Agent-Blackbox Cuts Claude Code Token Waste by 44%

Agent-Blackbox analyzes token usage for Claude Code and OpenCode. The tool reduces token waste by tracking local execution events. Optimizin

Recall Solves the Claude Code Cold Start Problem with Local Memory

Recall Solves the Claude Code Cold Start Problem with Local Memory

Recall provides a local memory layer for Claude Code to preserve session state. The tool uses TF-IDF and TextRank to summarize logs without

The Opaque ID Trick That Pushed Qwen 3:0.6B to 92% Accuracy

The Opaque ID Trick That Pushed Qwen 3:0.6B to 92% Accuracy

Qwen 3:0.6B achieved 92% classification accuracy through targeted fine-tuning. The team used Unsloth and QLora to optimize the small languag

ax-grep Cuts AI Agent Web Search Token Usage by 3x

ax-grep Cuts AI Agent Web Search Token Usage by 3x

ax-grep replicates browser accessibility trees to optimize web search. The tool reduces token usage by 3x and memory usage by 15x. It integr

The AI Coding Paradigm Shift: From Code Generation to Deterministic Harnesses

The AI Coding Paradigm Shift: From Code Generation to Deterministic Harnesses

AI coding is shifting from simple generation to a structured TDD pipeline. Deterministic harnesses are required to control LLM non-determini

Devin and the 80% Benchmark Rate That Commoditized AI Intelligence

Devin and the 80% Benchmark Rate That Commoditized AI Intelligence

AI agent benchmarks have surged from 13% to over 80% in 18 months. Code volume rose by 180% while actual deployments only grew by 30%. Marke

Why 13 Words Can Trick ChatGPT and Google AI Search

Why 13 Words Can Trick ChatGPT and Google AI Search

Cornell researchers found AI search agents are easily manipulated by short prompts. LLMs prioritize lexical similarity over factual accuracy

myDoo Launches VS Code Fork with Integrated Claude Code and Shared Memory

myDoo Launches VS Code Fork with Integrated Claude Code and Shared Memory

myDoo launched a VS Code-based IDE with Claude Code integrated. The platform features a RAG-based memory tool called Recall for multiple mod

Ponytail Cuts AI Code Volume by 54% to Stop Over-Engineering

Ponytail Cuts AI Code Volume by 54% to Stop Over-Engineering

Ponytail prevents AI coding agents from over-engineering software. Tests on FastAPI and React show a 54% reduction in code volume. The tool

Apertus: The Swiss Open Model Built for EU AI Act Compliance

Apertus: The Swiss Open Model Built for EU AI Act Compliance

The Swiss AI Initiative released Apertus as a fully transparent foundation model. It features 8B and 70B parameter versions supporting over

Why Anthropic is Now Requiring Government IDs for Claude

Why Anthropic is Now Requiring Government IDs for Claude

Anthropic now requires government IDs for certain Claude users. The verification process uses Persona to prevent platform abuse. This shift

Why NVIDIA Integrated OpenBao Into Its Serverless GPU Pipeline

Why NVIDIA Integrated OpenBao Into Its Serverless GPU Pipeline

NVIDIA officially adopted OpenBao for secret management in its cloud functions. The tool emerged as an open-source alternative to HashiCorp

Anthropic Fable 5 and Mythos 5 Hit by Trump Administration Export Controls

Anthropic Fable 5 and Mythos 5 Hit by Trump Administration Export Controls

The Trump administration ordered Anthropic to disable Fable 5 and Mythos 5. Amazon researchers triggered the move by bypassing model guardra

Why Apple Intelligence in iOS 27 Moves Beyond the Chatbot

Why Apple Intelligence in iOS 27 Moves Beyond the Chatbot

Apple integrates intelligence directly into core apps for iOS 27. The update shifts AI from conversational interfaces to embedded tools. On-