KO EN

AX BRIEF

AI news, benchmarks & engineering blog curation

● LIVE
지능1. Claude Fable 5.1 (max with fallback) 100 pt2. Claude Fable 5 (with fallback) 96 pt3. Claude Opus 5 (max) 95 pt4. GPT-6 Astra (max) 95 pt5. GPT-5.6 Sol (max) 91 pt코딩1. Claude Fable 5.1 (max) 100 pt2. Claude Opus 5 (xhigh) 94 pt3. GPT-6 Astra (max) 91 pt4. Muse Spark 1.3 (xhigh) 81 pt5. Grok 4.5 (high) 81 pt이미지1. GPT Image 2 (high) 100 pt2. MAI-Image-2.6 93 pt3. GPT Image 1.5 (high) 91 pt4. Reve 2.1 88 pt5. Muse Image 85 pt비디오1. Gemini Omni Flash 100 pt2. Minimax H3 Max (post-trained by fal) 94 pt3. Dreamina Seedance 2.0 720p 91 pt4. gemini-omni-1.1-flash 85 pt5. Wan 3.0 85 pt가격1. Claude Fable 5.1 (max with fallback) 100 pt2. GPT-6 Astra (max) 100 pt3. Claude Fable 5 (with fallback) 100 pt4. Claude Opus 5 (max) 49 pt5. GPT-5.6 Sol (max) 39 pt속도1. Gemini 3.5 Flash-Lite 100 pt2. Gemini 3.8 Flash (high) 80 pt3. Muse Spark 1.3 (max) 44 pt4. gpt-oss-120b (high) 40 pt5. GPT-5.6 Luna (max) 27 pt

AI Engineering Blogs

Curated deep insights from tech leaders and researchers — engineering, product, and strategy.

OpenAI Codex Security Framework: Controlling the Enterprise AI Agent

OpenAI Codex Security Framework: Controlling the Enterprise AI Agent

OpenAI revealed a security framework for its Codex AI agents. The system uses sandboxing and OpenTelemetry for strict governance. AI-driven

AMD ROCm Enables CUDA-Free Training for MedQA Medical AI

AMD ROCm Enables CUDA-Free Training for MedQA Medical AI

AMD ROCm successfully trained a MedQA model without using CUDA. The project utilized the AMD Instinct MI300X with 192GB of VRAM. High memory

Halliburton Accelerates Seismic Analysis Workflows by 95% via Amazon Bedrock

Halliburton Accelerates Seismic Analysis Workflows by 95% via Amazon Bedrock

Halliburton integrated Amazon Bedrock into its Seismic Engine software. The AI assistant automates the configuration of 82 specialized tools

OpenAI's GPT-Realtime-2 Boosts Audio Intelligence by 15.2%

OpenAI's GPT-Realtime-2 Boosts Audio Intelligence by 15.2%

OpenAI released three new real-time voice models via API. GPT-Realtime-2 shows a 15.2% improvement in audio intelligence. New pricing models

OpenAI Launches Three Realtime Voice Models and Transitions API to GA

OpenAI Launches Three Realtime Voice Models and Transitions API to GA

OpenAI released three specialized models for real-time voice applications. The GPT-Realtime-2 model features a 128K token context window for

Anthropic NLA Boosts Misalignment Detection From 3% to 15%

Anthropic NLA Boosts Misalignment Detection From 3% to 15%

Anthropic introduced Natural Language Autoencoders to decode AI activations. The system converts internal numerical states into human-readab

How Qwen2.5-0.5B and GRPO Eliminate LLM Reward Hacking

How Qwen2.5-0.5B and GRPO Eliminate LLM Reward Hacking

Amazon SageMaker AI implements RLVR and GRPO to stop reward hacking. Qwen2.5-0.5B improves math accuracy using the GSM8K dataset. Rule-based

AlphaEvolve Cuts Genomic Analysis Errors by 30 Percent

AlphaEvolve Cuts Genomic Analysis Errors by 30 Percent

Google's AlphaEvolve agent reduced genomic analysis errors by 30 percent. The AI optimizes the DeepConsensus model to improve mutation detec

The 5,000 Exaflop Scale of NVIDIA and DOE's Genesis Mission

The 5,000 Exaflop Scale of NVIDIA and DOE's Genesis Mission

The US Department of Energy and NVIDIA launched the Genesis Mission. The Solstice supercomputer will feature 100,000 Vera Rubin GPUs. AI is

Google DeepMind CEO Calls Korea an AI 'Best Restaurant' for Hardware

Google DeepMind CEO Calls Korea an AI 'Best Restaurant' for Hardware

Demis Hassabis praised Korea's manufacturing infrastructure as ideal for AI. Google's Gemini AI is now powering Boston Dynamics' Spot robot

vLLM V1 Transition: Solving Numerical Drift in RL Pipelines

vLLM V1 Transition: Solving Numerical Drift in RL Pipelines

vLLM transitioned from version 0.8.5 to 0.18.1 with major structural changes. Developers found numerical inconsistencies in GSPO reinforceme

KAIST and Google Use Physics-Informed AI to Map Geothermal Risks

KAIST and Google Use Physics-Informed AI to Map Geothermal Risks

KAIST received a Google Foundational Science Grant for geothermal research. The team uses physics-informed AI to predict underground fluid d

ZAYA1-8B Outperforms 119B Model in Math Using AMD Hardware

ZAYA1-8B Outperforms 119B Model in Math Using AMD Hardware

Zyphra AI released ZAYA1-8B trained on AMD Instinct MI300x hardware. The model uses 760M active parameters to beat much larger models in mat

Uber Deploys OpenAI Models to Optimize Driver Earnings and Voice Booking

Uber Deploys OpenAI Models to Optimize Driver Earnings and Voice Booking

Uber integrated OpenAI models to help drivers optimize their daily earnings. A multi-agent architecture and AI Guard ensure low latency and

Parloa AMP Scales Enterprise Voice AI Using Modular GPT-5.4 Agents

Parloa AMP Scales Enterprise Voice AI Using Modular GPT-5.4 Agents

Parloa launched the AMP platform to manage enterprise voice AI agents. The system uses modular GPT-5.4 agents to reduce human transfers by 8

OpenAI's MRC Protocol Solves the AI Supercomputer Network Bottleneck

OpenAI's MRC Protocol Solves the AI Supercomputer Network Bottleneck

OpenAI introduced the MRC protocol to eliminate GPU idle time in large clusters. The system uses intelligent packet-spraying to connect 130,

NeuralBench: Meta AI Unifies 94 Datasets for Brain-Wave AI

NeuralBench: Meta AI Unifies 94 Datasets for Brain-Wave AI

Meta AI released NeuralBench to standardize EEG model evaluation. The framework integrates 94 datasets and 36 distinct brain tasks. Results

OpenAI B2B Signals Reveal a 3.5x AI Intelligence Gap in Enterprises

OpenAI B2B Signals Reveal a 3.5x AI Intelligence Gap in Enterprises

OpenAI B2B Signals show leading firms use AI 3.5x more than peers. Cisco saved 1,500 engineering hours monthly using Codex. Enterprise AI su

Anthropic Partners With SpaceX to Double Claude Code Usage Limits

Anthropic Partners With SpaceX to Double Claude Code Usage Limits

Anthropic is utilizing SpaceX Colossus 1 resources to expand compute capacity. Claude Code usage limits have doubled for Pro, Max, Team, and

NVIDIA Spectrum-X and MRC: Solving the Giga-scale AI Bottleneck

NVIDIA Spectrum-X and MRC: Solving the Giga-scale AI Bottleneck

NVIDIA introduced the MRC protocol to Spectrum-X for AI networks. The protocol enables multipath RDMA to reduce GPU idle time. The technolog

Open ASR Leaderboard Introduces Private Datasets to Combat Data Leakage

Open ASR Leaderboard Introduces Private Datasets to Combat Data Leakage

The Open ASR Leaderboard added private datasets to prevent benchmark contamination. New evaluation standards use OpenAI Whisper's normalizat

OpenAI's MRC Protocol: The Architecture Linking 131,000 GPUs

OpenAI's MRC Protocol: The Architecture Linking 131,000 GPUs

OpenAI released the MRC protocol to the Open Compute Project. The system connects 131,000 GPUs using only two switch layers. MRC reduces tra

The 3x Inference Boost Powering Google Gemma 4

The 3x Inference Boost Powering Google Gemma 4

Google introduced MTP drafters to accelerate Gemma 4 inference. The technology increases generation speed by up to 3x without quality loss.

OpenAI Launches ChatGPT Ads Manager with CPC Bidding for US Brands

OpenAI Launches ChatGPT Ads Manager with CPC Bidding for US Brands

OpenAI released a beta Ads Manager for US-based companies to run campaigns. The platform now supports CPC bidding alongside traditional CPM

Project Arc: NVIDIA and ServiceNow Shift AI From Chatbots to Autonomous Agents

Project Arc: NVIDIA and ServiceNow Shift AI From Chatbots to Autonomous Agents

NVIDIA and ServiceNow unveiled Project Arc for autonomous enterprise tasks. The system leverages Blackwell GPUs to reduce token costs by 35

PORTool Uses Reward Trees to Optimize AI Agent Tool Calling

PORTool Uses Reward Trees to Optimize AI Agent Tool Calling

PORTool introduces a reward tree to optimize AI agent tool calling. The algorithm reduces redundant steps while increasing final accuracy. T

The 64.37% Benchmark Driving Anthropic's 10 New Financial Agents

The 64.37% Benchmark Driving Anthropic's 10 New Financial Agents

Anthropic released 10 specialized agent templates for financial services. Claude Opus 4.7 achieves 64.37% accuracy on the Vals AI benchmark.

Gemini API Webhooks Replace Polling for Long-Running Operations

Gemini API Webhooks Replace Polling for Long-Running Operations

Google introduced event-driven webhooks to the Gemini API. The update eliminates inefficient polling for long-running AI tasks. Developers c

Amazon Quick S3 Tables Integration Eliminates Data Pipeline Latency

Amazon Quick S3 Tables Integration Eliminates Data Pipeline Latency

Amazon Quick now integrates directly with S3 Tables. The update removes the need for complex ETL pipelines. Users can analyze Apache Iceberg

OpenAI and PwC Partner to Deploy AI Agents for Financial Automation

OpenAI and PwC Partner to Deploy AI Agents for Financial Automation

OpenAI and PwC are collaborating to build AI agents for corporate finance. OpenAI processed five times more contracts using its own Codex mo