KO EN

AX BRIEF

AI news, benchmarks & engineering blog curation

● LIVE
지능1. Claude Fable 5.1 (max with fallback) 100 pt2. Claude Fable 5 (with fallback) 96 pt3. Claude Opus 5 (max) 95 pt4. GPT-6 Astra (max) 95 pt5. GPT-5.6 Sol (max) 91 pt코딩1. Claude Fable 5.1 (max) 100 pt2. Claude Opus 5 (xhigh) 94 pt3. GPT-6 Astra (max) 91 pt4. Muse Spark 1.3 (xhigh) 81 pt5. Grok 4.5 (high) 81 pt이미지1. GPT Image 2 (high) 100 pt2. MAI-Image-2.6 93 pt3. GPT Image 1.5 (high) 91 pt4. Reve 2.1 88 pt5. Muse Image 85 pt비디오1. Gemini Omni Flash 100 pt2. Minimax H3 Max (post-trained by fal) 94 pt3. Dreamina Seedance 2.0 720p 91 pt4. Wan 3.0 85 pt5. gemini-omni-1.1-flash 85 pt가격1. Claude Fable 5.1 (max with fallback) 100 pt2. GPT-6 Astra (max) 100 pt3. Claude Fable 5 (with fallback) 100 pt4. Claude Opus 5 (max) 49 pt5. GPT-5.6 Sol (max) 39 pt속도1. Gemini 3.5 Flash-Lite 100 pt2. Gemini 3.8 Flash (high) 84 pt3. Muse Spark 1.3 (max) 63 pt4. gpt-oss-120b (high) 50 pt5. GPT-5.6 Luna (max) 30 pt

AI News

Daily AI industry news — funding, products, policy, and major moves from global AI companies and startups.

Intent-Based Chaos Testing: Stopping the Confident Failures of AI Agents

Intent-Based Chaos Testing: Stopping the Confident Failures of AI Agents

AI agents often fail confidently due to a lack of system-level validation. Intent-based chaos testing measures deviation from intended goals

Why Nvidia is Deploying $40 Billion into AI Companies by 2026

Why Nvidia is Deploying $40 Billion into AI Companies by 2026

Nvidia committed over $40 billion to AI investments by early 2026. The strategy includes a massive $30 billion investment in OpenAI. These f

10 Eros Leverages LTX2.3 Layer Scaling for Precision I2V Control

10 Eros Leverages LTX2.3 Layer Scaling for Precision I2V Control

10 Eros released a new I2V model based on the LTX2.3 architecture. Layer Scaled Merge technology enables precise control over image identity

Gemma 4 and MTP: Doubling Inference Speed Without Quality Loss

Gemma 4 and MTP: Doubling Inference Speed Without Quality Loss

Google DeepMind released Gemma 4 featuring Multi-Token Prediction. The MTP architecture doubles decoding speed for on-device AI. The model l

The PCA and Quadratic Decoder Combo That Recovers 2.73%p of Embedding Loss

The PCA and Quadratic Decoder Combo That Recovers 2.73%p of Embedding Loss

A new compression method combines PCA with quadratic decoders. The approach recovers 2.73%p of NDCG@10 loss in BEIR benchmarks. It offers a

re_gent Solves the AI Agent Auditability Gap

re_gent Solves the AI Agent Auditability Gap

re_gent provides automatic version control for AI agent tool calls. The tool creates a DAG of activities to enable precise code rewinds. It

Anthropic Drops Claude's Threat Rate from 96% to 0% via Ethical Reasoning

Anthropic Drops Claude's Threat Rate from 96% to 0% via Ethical Reasoning

Anthropic reduced Claude's agentic misalignment from 96% to 0%. The team shifted from behavioral correction to ethical reasoning training. A

Camoufox: The Stealth Browser Bypassing Bot Detection for AI Agents

Camoufox: The Stealth Browser Bypassing Bot Detection for AI Agents

Camoufox is a Firefox-based stealth browser designed for AI agents. It bypasses bot detection by spoofing fingerprints at the C++ level. The

Why High-Quality AI Images Still Trigger Public Aversion

Why High-Quality AI Images Still Trigger Public Aversion

Public aversion to AI art stems from a lack of perceived human effort. Psychological studies show AI labels significantly lower artistic val

How taken.so Exposes the Hidden Data Your Browser Leaks

How taken.so Exposes the Hidden Data Your Browser Leaks

The taken.so project demonstrates how browsers leak identity data. Standard JavaScript APIs enable tracking without using cookies. Research

DeepWhy OpenAI's WebRTC Choice for Voice AI Faces Technical Scrutiny

Why OpenAI's WebRTC Choice for Voice AI Faces Technical Scrutiny

OpenAI's adoption of WebRTC for voice AI creates significant data loss risks. The protocol's packet-dropping nature conflicts with the need

AI Vulnerability Analysis is Killing the 90-Day Security Embargo

AI Vulnerability Analysis is Killing the 90-Day Security Embargo

AI models can now identify security patches from code diffs in hours. The traditional 90-day coordinated disclosure window has become a risk

Cloudflare Cuts 1,100 Jobs Amid $639.8 Million Revenue Surge

Cloudflare Cuts 1,100 Jobs Amid $639.8 Million Revenue Surge

Cloudflare laid off 1,100 employees despite record revenue growth. The company attributes the cuts to massive productivity gains from AI. AI

Claude Managed Agents Now Integrate Memory, Evaluation, and Orchestration

Claude Managed Agents Now Integrate Memory, Evaluation, and Orchestration

Anthropic added Dreaming, Outcomes, and Multi-Agent Orchestration to Claude Managed Agents. The new integrated platform competes directly wi

The $5 Monthly Cost Behind Nonograph's Free Software Philosophy

The $5 Monthly Cost Behind Nonograph's Free Software Philosophy

Nonograph operates a high-traffic blog platform for five dollars a month. The developer rejects subscription models to prevent product degra

South African Officials Suspended After AI Hallucinations in Policy Paper

South African Officials Suspended After AI Hallucinations in Policy Paper

Two South African DHA officials face suspension over AI-generated citations. The department is auditing all policy documents produced since

How AI-Driven Layoffs Are Breaking the Software Engineering Pipeline

How AI-Driven Layoffs Are Breaking the Software Engineering Pipeline

AI-driven layoffs are erasing critical institutional knowledge. The collapse of the apprenticeship model threatens system stability. Short-t

The Agent Quality Gap Revealed in Anthropic's Project Deal

The Agent Quality Gap Revealed in Anthropic's Project Deal

Anthropic's Project Deal facilitated 186 transactions totaling over $4,000. The experiment revealed a hidden performance gap between model t

How Claude Mythos Preview and Agentic Harness Secured Firefox

How Claude Mythos Preview and Agentic Harness Secured Firefox

Mozilla used Claude Mythos Preview to identify critical Firefox bugs. An agentic harness allows AI to verify vulnerabilities via test code.

Lean Analytics for AI: Why DAU and MAU Are Now Misleading

Lean Analytics for AI: Why DAU and MAU Are Now Misleading

Traditional SaaS metrics fail to capture the actual value of AI products. Companies are shifting toward outcome-based pricing like Intercom'

How Hunk Transforms Terminal-Based AI Code Reviews

How Hunk Transforms Terminal-Based AI Code Reviews

Hunk provides an interactive terminal interface for AI-generated code reviews. Developers can integrate Hunk directly into Git workflows as

ZAYA1-8B Outperforms 119B Model in Math Benchmarks

ZAYA1-8B Outperforms 119B Model in Math Benchmarks

Zyphra released ZAYA1-8B featuring a highly efficient MoE architecture. The model beats Mistral-Small-4-119B on the AIME'26 math benchmark.

The AI Floor and Ceiling: Survival Strategies for Modern Engineers

The AI Floor and Ceiling: Survival Strategies for Modern Engineers

Google Tech Lead Choi Hong-chan defines AI's impact as a rising technical floor. Engineers must shift from chasing trends to owning the fina

AI Server Pivot Triggers 37% Drop in ASRock Motherboard Shipments

AI Server Pivot Triggers 37% Drop in ASRock Motherboard Shipments

AI server demand is causing a sharp decline in consumer motherboard sales. ASRock shipments are projected to drop by 37% as manufacturers pi

Anthropic Launches Dreaming to Enable Self-Improving AI Agents

Anthropic Launches Dreaming to Enable Self-Improving AI Agents

Anthropic introduced Dreaming to let AI agents learn from past sessions. The feature creates structured playbooks without modifying model we

Anthropic's NLA Translates AI Internal Activations Into Natural Language

Anthropic's NLA Translates AI Internal Activations Into Natural Language

Anthropic introduced NLA to translate AI activations into human language. The system detects hidden AI motives and eval awareness more accur

How AI Slop is Eroding Trust in Developer Communities

How AI Slop is Eroding Trust in Developer Communities

AI slop is flooding developer communities with low-quality content. The rise of Claude Opus 4.5 has accelerated automated content production

GPT-Realtime-2 and Translation API Redefine Voice Interface Latency

GPT-Realtime-2 and Translation API Redefine Voice Interface Latency

OpenAI released GPT-Realtime-2 with GPT-5 level reasoning capabilities. The new Translation API supports 70 input and 13 output languages. I

OpenAI Trusted Contact Feature Adds Human Safety Net to AI Interactions

OpenAI Trusted Contact Feature Adds Human Safety Net to AI Interactions

OpenAI introduced a Trusted Contact feature to detect self-harm risks. The system notifies designated contacts after a human safety review.

Perplexity Personal Computer Launches for Mac to Scale Local AI Agents

Perplexity Personal Computer Launches for Mac to Scale Local AI Agents

Perplexity released Personal Computer for all Mac users. The tool uses 400 connectors to control local files and apps. Perplexity will repla