KO EN

AX BRIEF

AI news, benchmarks & engineering blog curation

● LIVE
지능1. Claude Fable 5.1 (max with fallback) 100 pt2. Claude Fable 5 (with fallback) 96 pt3. Claude Opus 5 (max) 95 pt4. GPT-6 Astra (max) 95 pt5. GPT-5.6 Sol (max) 91 pt코딩1. Claude Fable 5.1 (max) 100 pt2. Claude Opus 5 (xhigh) 94 pt3. GPT-6 Astra (max) 91 pt4. Muse Spark 1.3 (xhigh) 81 pt5. Grok 4.5 (high) 81 pt이미지1. GPT Image 2 (high) 100 pt2. MAI-Image-2.6 93 pt3. GPT Image 1.5 (high) 91 pt4. Reve 2.1 88 pt5. Muse Image 85 pt비디오1. Gemini Omni Flash 100 pt2. Minimax H3 Max (post-trained by fal) 94 pt3. Dreamina Seedance 2.0 720p 91 pt4. Wan 3.0 85 pt5. gemini-omni-1.1-flash 85 pt가격1. Claude Fable 5.1 (max with fallback) 100 pt2. GPT-6 Astra (max) 100 pt3. Claude Fable 5 (with fallback) 100 pt4. Claude Opus 5 (max) 49 pt5. GPT-5.6 Sol (max) 39 pt속도1. Gemini 3.5 Flash-Lite 100 pt2. Muse Spark 1.3 (max) 44 pt3. gpt-oss-120b (high) 36 pt4. GPT-5.6 Luna (max) 28 pt5. GPT-5.6 Terra (max) 21 pt

AI Engineering Blogs

Curated deep insights from tech leaders and researchers — engineering, product, and strategy.

Gemini 3.5 Transcribe Cuts Transcription Time by 70%

Gemini 3.5 Transcribe Cuts Transcription Time by 70%

Google released Gemini 3.5 Transcribe to convert raw audio into refined text. The model reduces transcription time by 70% compared to Chirp

Why mLateOn-medical Outperforms General Models in Medical Retrieval

Why mLateOn-medical Outperforms General Models in Medical Retrieval

mLateOn-medical outperforms general models in medical retrieval. The model was trained on one RTX 3090 GPU for 14.5 hours. Multi-vector arch

How loveholidays Moved 79% of Code Deployments to Non-Developers via Codex

How loveholidays Moved 79% of Code Deployments to Non-Developers via Codex

loveholidays increased AI-supported code changes from 7% to 79% in one year. Non-developers now deploy features using a Codex-powered Search

The Python Data Class Config That Slashes Memory Overhead

The Python Data Class Config That Slashes Memory Overhead

Python data classes eliminate repetitive boilerplate code for developers. The slots=True setting significantly reduces memory overhead in la

Amazon Quick Desktop and FSx for ONTAP Cut Reporting from Hours to Minutes

Amazon Quick Desktop and FSx for ONTAP Cut Reporting from Hours to Minutes

Amazon Quick Desktop automates weekly business reporting through skill-based workflows. FSx for ONTAP and S3 Access Points ensure strict da

Amazon OpenSearch MCP Apps Close the Observability Verification Gap

Amazon OpenSearch MCP Apps Close the Observability Verification Gap

Amazon OpenSearch MCP Apps integrate visual data directly into AI chat. The dual response pattern eliminates the manual verification gap for

OpenAI Jalapeño: The Inference Chip Reducing Latency by 3.6x

OpenAI Jalapeño: The Inference Chip Reducing Latency by 3.6x

OpenAI unveiled Jalapeño, a custom inference chip reducing latency by 3.6x. The hardware achieves 1.9x better power efficiency for large-sca

OpenAI's Jalapeño Chip and the Quest for Useful Intelligence per Dollar

OpenAI's Jalapeño Chip and the Quest for Useful Intelligence per Dollar

OpenAI unveiled Jalapeño, its first in-house inference chip for efficiency. The chip integrates with a full-stack co-design strategy and new

NVIDIA RTX Spark Integrates AI Agents and Gaming on a Single Chip

NVIDIA RTX Spark Integrates AI Agents and Gaming on a Single Chip

NVIDIA announced the RTX Spark platform for Windows PCs this autumn. The system integrates AI agents and gaming on a single-chip architectur

Google Search AI Now Maps Furniture to Your Actual Room Dimensions

Google Search AI Now Maps Furniture to Your Actual Room Dimensions

Google Search AI introduces five new tools for home decoration. Visual context and real-time video now drive furniture discovery. Integrated

ChatGPT Work and Codex Admin Plugin Automates Workspace Governance

ChatGPT Work and Codex Admin Plugin Automates Workspace Governance

OpenAI released an Admin Plugin for ChatGPT Work and Codex. The tool enables workspace management through natural language. Administrators c

IBM Granite 4.2 Shifts From Instruction Following to Explicit Reasoning

IBM Granite 4.2 Shifts From Instruction Following to Explicit Reasoning

IBM released Granite 4.2 with 3B, 8B, and 30B model sizes. The models feature a 512K context window and 15 trillion tokens of training. Asyn

Granite Speech 5.0 Turbo CTC Hits 12,600 RTFx Processing Speed

Granite Speech 5.0 Turbo CTC Hits 12,600 RTFx Processing Speed

IBM released Granite Speech 5.0 Turbo CTC with 12,600 RTFx speed. The encoder-only architecture enables 3.5 hours of transcription per secon

Why QAH Direct Distillation Beats the 16-Bit Performance Ceiling

Why QAH Direct Distillation Beats the 16-Bit Performance Ceiling

QAH allows 4-bit models to outperform their 16-bit originals. Direct distillation from full-size models removes performance ceilings. The me

How a Translation Error Exposed Russia's ChatGPT Influence Operation

How a Translation Error Exposed Russia's ChatGPT Influence Operation

OpenAI banned a Russian influence network using ChatGPT accounts. The group created a fake Israeli institute to spread propaganda. A transla

Amazon Bedrock Automates the Metadata Bottleneck in Data Engineering

Amazon Bedrock Automates the Metadata Bottleneck in Data Engineering

Amazon Bedrock powers a new automated metadata correction system. The system uses a hierarchical approach to minimize LLM inference costs. H

gr.Workflow Collapses AI Pipeline Design Into Instant REST APIs

gr.Workflow Collapses AI Pipeline Design Into Instant REST APIs

Gradio introduced gr.Workflow to design AI pipelines via a node-based canvas. The visual graph automatically generates both a user interface

Amazon Connect and MCP Enable App-Free AI Phone Ordering

Amazon Connect and MCP Enable App-Free AI Phone Ordering

Amazon Connect and MCP enable AI phone ordering without apps or logins. The architecture uses Claude Haiku 4.5 and a decoupled backend via M

DeepAWS Knowledge Management Accelerator Preserves Retiring Expertise via Voice Avatars

AWS Knowledge Management Accelerator Preserves Retiring Expertise via Voice Avatars

AWS launched a knowledge management accelerator that captures retiring workers' expertise through voice-enabled avatars. The system uses Ama

DeepVera Rubin NVL72 Boosts Throughput per Megawatt 30x for Agent AI

Vera Rubin NVL72 Boosts Throughput per Megawatt 30x for Agent AI

NVIDIA Vera Rubin NVL72 increases throughput per megawatt by 30x. The system reduces agent AI token costs by up to 35x over GB300. New hardw

Why SageMaker HyperPod's Ray Integration Removes the Need for kubectl

Why SageMaker HyperPod's Ray Integration Removes the Need for kubectl

AWS integrated Ray into SageMaker HyperPod to simplify cluster management. The update removes the need for manual YAML manifests and kubectl

Why ARD is the New DNS for Multi-Cloud AI Agents

Why ARD is the New DNS for Multi-Cloud AI Agents

AWS Agent Registry and the ARD open standard solve AI resource silos. ARD uses a DNS-like federation to discover agents across multi-cloud e

NVIDIA Groq 3 LPX Solves the Decode Latency Bottleneck for Agent AI

NVIDIA Groq 3 LPX Solves the Decode Latency Bottleneck for Agent AI

NVIDIA Groq 3 LPX achieves 3,400 tokens per second in long-context environments. The system splits workloads between Rubin GPUs and LPUs to

DeepWhy Local SLMs and Ollama are Replacing Cloud APIs for Devs

Why Local SLMs and Ollama are Replacing Cloud APIs for Devs

Local SLMs offer a secure alternative to cloud APIs by keeping data on-device. Ollama simplifies the deployment of quantized models across v

Grok Build Collapsed Data Analysis and API Deployment into 4 Prompts

Grok Build Collapsed Data Analysis and API Deployment into 4 Prompts

xAI released Grok Build as a TUI-based coding agent powered by Grok 4.6. The tool automates the entire ML pipeline from data cleaning to clo

DeepEPFL's 150-Microgram Robot Turns Sound Vibrations Into Thrust

EPFL's 150-Microgram Robot Turns Sound Vibrations Into Thrust

EPFL developed a 150-microgram robot powered by sound vibrations. The device uses Helmholtz resonance to create active directional thrust. T

Foundation World Models: Trading Average Accuracy for Trust Guarantees

Foundation World Models: Trading Average Accuracy for Trust Guarantees

Florent Delgrange proposes a new framework for autonomous agent world models. The system prioritizes fitness for guarantee over simple predi

DeepAmazon Bedrock RAG: Cutting Token Costs by 33% via Query-Aware Compression

Amazon Bedrock RAG: Cutting Token Costs by 33% via Query-Aware Compression

Amazon Bedrock RAG costs drop by 33% using query-aware compression. Claude Haiku filters noise to reduce input tokens by up to 10.1 times. H

AI Agents Shift From Answering to Executing Across 5 Key Industries

AI Agents Shift From Answering to Executing Across 5 Key Industries

AI agents are transitioning from text generation to autonomous execution. Five industries are seeing response times drop from days to minute

Muse Glimmer Hits 127 Tokens Per Second for Local Autonomous Coding

Muse Glimmer Hits 127 Tokens Per Second for Local Autonomous Coding

Muse Glimmer enables autonomous coding on local hardware. Speculative decoding pushes inference speeds to 127 tokens per second. The model e