AI news, benchmarks & engineering blog curation
Curated deep insights from tech leaders and researchers — engineering, product, and strategy.

Google released Gemini 3.5 Transcribe to convert raw audio into refined text. The model reduces transcription time by 70% compared to Chirp

mLateOn-medical outperforms general models in medical retrieval. The model was trained on one RTX 3090 GPU for 14.5 hours. Multi-vector arch

loveholidays increased AI-supported code changes from 7% to 79% in one year. Non-developers now deploy features using a Codex-powered Search

Python data classes eliminate repetitive boilerplate code for developers. The slots=True setting significantly reduces memory overhead in la

Amazon Quick Desktop automates weekly business reporting through skill-based workflows. FSx for ONTAP and S3 Access Points ensure strict da

Amazon OpenSearch MCP Apps integrate visual data directly into AI chat. The dual response pattern eliminates the manual verification gap for

OpenAI unveiled Jalapeño, a custom inference chip reducing latency by 3.6x. The hardware achieves 1.9x better power efficiency for large-sca

OpenAI unveiled Jalapeño, its first in-house inference chip for efficiency. The chip integrates with a full-stack co-design strategy and new

NVIDIA announced the RTX Spark platform for Windows PCs this autumn. The system integrates AI agents and gaming on a single-chip architectur

Google Search AI introduces five new tools for home decoration. Visual context and real-time video now drive furniture discovery. Integrated

OpenAI released an Admin Plugin for ChatGPT Work and Codex. The tool enables workspace management through natural language. Administrators c

IBM released Granite 4.2 with 3B, 8B, and 30B model sizes. The models feature a 512K context window and 15 trillion tokens of training. Asyn

IBM released Granite Speech 5.0 Turbo CTC with 12,600 RTFx speed. The encoder-only architecture enables 3.5 hours of transcription per secon

QAH allows 4-bit models to outperform their 16-bit originals. Direct distillation from full-size models removes performance ceilings. The me

OpenAI banned a Russian influence network using ChatGPT accounts. The group created a fake Israeli institute to spread propaganda. A transla

Amazon Bedrock powers a new automated metadata correction system. The system uses a hierarchical approach to minimize LLM inference costs. H

Gradio introduced gr.Workflow to design AI pipelines via a node-based canvas. The visual graph automatically generates both a user interface

Amazon Connect and MCP enable AI phone ordering without apps or logins. The architecture uses Claude Haiku 4.5 and a decoupled backend via M

AWS launched a knowledge management accelerator that captures retiring workers' expertise through voice-enabled avatars. The system uses Ama

NVIDIA Vera Rubin NVL72 increases throughput per megawatt by 30x. The system reduces agent AI token costs by up to 35x over GB300. New hardw

AWS integrated Ray into SageMaker HyperPod to simplify cluster management. The update removes the need for manual YAML manifests and kubectl

AWS Agent Registry and the ARD open standard solve AI resource silos. ARD uses a DNS-like federation to discover agents across multi-cloud e

NVIDIA Groq 3 LPX achieves 3,400 tokens per second in long-context environments. The system splits workloads between Rubin GPUs and LPUs to

Local SLMs offer a secure alternative to cloud APIs by keeping data on-device. Ollama simplifies the deployment of quantized models across v

xAI released Grok Build as a TUI-based coding agent powered by Grok 4.6. The tool automates the entire ML pipeline from data cleaning to clo

EPFL developed a 150-microgram robot powered by sound vibrations. The device uses Helmholtz resonance to create active directional thrust. T

Florent Delgrange proposes a new framework for autonomous agent world models. The system prioritizes fitness for guarantee over simple predi

Amazon Bedrock RAG costs drop by 33% using query-aware compression. Claude Haiku filters noise to reduce input tokens by up to 10.1 times. H

AI agents are transitioning from text generation to autonomous execution. Five industries are seeing response times drop from days to minute

Muse Glimmer enables autonomous coding on local hardware. Speculative decoding pushes inference speeds to 127 tokens per second. The model e