Today's AI updates span major model scaling, image generation upgrades, and developer tool enhancements across the industry. OpenAI's Bell model scales to 10 trillion parameters utilizing a dual-speed learning system and the Recursive Self-Improvement (RSI) index to enable continuous evolution, while OpenAI also acknowledges that deidentified product usage data and researcher data may have contributed to its model improvements and scientific solutions. Meanwhile, ChatGPT has introduced image editing tools that enable targeted modifications using a comment-based system, and the GPT image 2.5 update introduces significant performance gains in generation speed alongside that same comment-based editing system for enhanced precision. In visual reasoning, a new internal Gemini Pro checkpoint demonstrates high-reasoning capabilities through the generation of detailed SVG graphics. Furthermore, DeepSeek V4.1 Flash launches with open weights to provide competitive performance against frontier models, featuring a controllable reasoning effort parameter and support for private local deployment via quantization. Hardware and software integrations also advance, as Apple announces two new Apple Watches featuring ambient listening capabilities for local processing, and Codex expands its context window to 828,000 tokens via configuration files despite maintaining strict limits on skill lists.
01OpenAI's Bell Model Scales to 10 Trillion Parameters
OpenAI is attempting to solve one of the biggest limitations of modern AI: the fact that models typically stop learning the moment they are released to the public. The company has reportedly completed a pre-training run for a model codenamed Bell, which scales to more than 10 trillion parameters. To put that in perspective, GPT4 sat in the trillion range. Rather than remaining a static tool, Bell is designed for continuous evolution, moving the industry toward models that can grow and refine their own intelligence after launch.
This evolution is powered by a dual-speed learning system. While traditional models are "frozen" after their initial training, Bell utilizes a fast weight layer to immediately soak up lessons from verified proofs, experiments, code tests, and tool traces. This allows the model to learn from its immediate environment in real time. A secondary, slower loop then filters these experiences, consolidating only the most effective lessons into persistent weights and training recipes. This two-tiered approach allows the model to balance rapid adaptation with long-term stability.
To measure this capability, OpenAI developed the Recursive Self-Improvement (RSI) index. This index tracks a model's ability to evolve by aggregating metrics from machine learning experiments, research debugging, and the optimization of kernels and training recipes. The practical application of this is seen in the Saul model, which outperformed GPT 5.5 by 16.2 points on the RSI index. Saul effectively acted as its own engineer, designing hundreds of architecture experiments for a smaller draft model and managing its training. Human intervention was only required for training instability or hardware failures, and the process resulted in a token generation efficiency increase of more than 15%.
02ChatGPT has introduced image editing tools that enable targe
Users can now modify specific parts of an AI-generated image without the frustration of regenerating the entire picture from scratch. ChatGPT has introduced a system that effectively transforms static images into editable canvases, allowing for targeted modifications. Instead of relying on a new prompt and hoping for a similar result, users can pinpoint a specific element—such as a person's clothing or a piece of text—and request a precise change. This shift moves the AI image workflow away from a trial-and-error process and toward a controlled editing environment where the user has direct agency over the final composition.
The technical implementation relies on a comment-based system for directing the AI. While Google Pix utilizes direct element isolation to facilitate changes, ChatGPT allows users to select an element and then use a comment to specify exactly what should be altered. For instance, the tool is capable of identifying distinct components within a scene, such as a man wearing earbuds, the earbuds' charging case, or the wireless earbuds themselves. A user could select a specific piece of text, such as a product rating of 4.9 out of five, and simply comment that it should be changed to 4.8 to feel more accurate. This allows for granular control over colors, subjects, and text without disturbing the rest of the image.
To further streamline the creative workflow, ChatGPT now supports the batching of multiple targeted edits. Rather than applying changes one by one and waiting for the image to refresh after every single tweak, users can queue several specific modifications. This means a user could simultaneously swap a subject's gender, change the color of an object, and update text before applying all the edits to the final image in a single step. This batching capability, which is also a feature of Google Pix, ensures that the final output is a result of a curated set of changes rather than a series of disjointed iterations. By saving multiple edits before sending the final request, users can refine complex images with much greater efficiency.
03There is no official technical documentation or performance
The mystery surrounding the rumored BEL model from OpenAI leaves the industry in a state of speculation because there is no verifiable way to measure its capabilities or integrate it into existing workflows. For developers and businesses, the absence of technical specifications means they cannot plan for its adoption or compare it against existing tools. This lack of transparency transforms what could be a major technological leap into a series of unverified rumors, as there is no concrete proof that the model exists in a usable or standardized form.
A detailed breakdown conducted on September 6th by Marcus Bell reveals just how sparse the available information is. There has been no official announcement from OpenAI, nor is there a model card, which is the standard document that outlines a model's training data and intended purpose. Furthermore, there are no benchmarks, the standardized performance tests used to rank AI models against one another. The missing data extends to the practicalities of deployment: there is no public pricing, no API for software integration, no released weights that would allow others to run the model, and no specified license governing its use.
Even the most fundamental characteristics of the model remain unknown. There is no confirmed context window, which refers to the maximum amount of information the model can consider at one time during a conversation. Perhaps most strikingly, the confirmed modality—whether the model processes text, images, or other types of data—has not been established. While OpenAI has made references to an internal model, the company has notably avoided using the name BEL in its statements. This gap between the rumors and the official record suggests that the evidence for the model's existence and performance is currently thin, leaving the community with no official technical foundation to rely upon.
04OpenAI Admits Use of Deidentified Product Data
Your interactions with an AI tool, even when stripped of personal identifiers, may be contributing to the very intelligence the tool provides. OpenAI recently acknowledged that it cannot rule out the possibility that deidentified product usage data has been used to improve its models. Deidentified data refers to information where personal details like names or email addresses are removed, but the core content of the interaction remains. This admission suggests that the logic, problem-solving steps, and patterns provided by users can still be absorbed into the system to enhance its overall capabilities and scientific solutions.
This issue became prominent after mathematicians Tristan Buckmaster and Levent Alpaga raised concerns regarding the timeline and data usage surrounding OpenAI's approach to the Navier Stokes problem. The researchers questioned how the company arrived at its solutions, leading OpenAI to admit that it could not rule out that deidentified data derived from the researchers' own usage of the products helped improve the models. While OpenAI characterized this outcome as unlikely, the acknowledgment confirms that the intellectual labor of high-level users can potentially be repurposed to refine the AI's performance on specialized tasks.
For professional users and scientific researchers, this creates a significant tension regarding intellectual property and privacy. While deidentification protects a person's identity, it does not necessarily protect the unique methodology or the proprietary breakthroughs they develop while using the software. When a model learns from the way a mathematician structures a proof or a scientist approaches a formula, the AI is effectively capturing the expertise of the user. This suggests that the boundary between a user's private work environment and the model's training data is more porous than previously understood, as the patterns of usage themselves become a valuable resource for model improvement.
05GPT-6 Astra Integrates Hardware and Figma
GPT 6 Astra is evolving from a chatbot into a functional AI employee capable of managing complex software and physical hardware. Rather than simply answering prompts, Astra can operate a computer autonomously, navigating through applications to complete multi-step tasks. This includes performing creative work directly within third-party software like Figma and Blender, or even interacting with robots. To ensure these autonomous actions remain reliable, users can implement a "stop and ask" rule. This constraint forces the AI to pause and seek human confirmation rather than guessing when it encounters an ambiguity, ensuring that the human retains final decision-making authority in a hybrid workflow.
The operational scope of Astra extends into business management, where it can function as a Chief Marketing Officer. In this capacity, the model does not just create content; it manages the entire deployment and optimization cycle. Astra can push finished marketing materials live, monitor their real-world performance, and autonomously rebuild specific elements that are not resonating with the audience. By handling this entire loop, Astra removes the need for external teams to manage the iterative process of performance tuning.
While Astra expands utility, OpenAI is developing an even more powerful internal checkpoint known as Bell. This next-generation model is significantly more capable, with reports indicating that Bell's lowest reasoning level already outperforms GPT 6 Astra at its maximum reasoning capacity. The scale of Bell's intelligence was recently demonstrated when OpenAI utilized a massive multi-agent system of 10,000 Bell-powered agents to propose a solution for the Navier-Stokes Millennium Prize problem. This experiment involved 88 hours of work and the exchange of 2.7 million messages, generating 130 billion output tokens within a $1 million budget.
Parallel to these developments, DeepSeek V4.1 Flash is introducing extreme cost efficiency and visual autonomy. The model can autonomously generate landing pages—such as for mehmoodmohammad.com—using Command Code's design skill harness. It employs a visual verification loop, using an agent browser to take screenshots and a scratch pad to verify that images are properly cropped or zoomed. This efficiency is reflected in the cost; one session to transform a landing page cost only 7 cents, and a high-volume usage of 3 billion tokens cost approximately $30. Furthermore, V4.1 Flash shows superior generalization in its post-training reinforcement learning, maintaining high performance across various versions of the Terminal Bench.
06Gemini Pro Checkpoint Generates Complex SVGs
Google has released a new internal version of its AI, known as a Gemini Pro Checkpoint, which suggests a significant leap in the model's ability to handle complex, multi-step reasoning tasks. This update, which may be an early glimpse of the upcoming Gemini 4.0 Pro model series, represents the first new Pro Checkpoint seen in a considerable amount of time. For users and developers, this indicates that Google is pushing the boundaries of how AI can translate abstract concepts into precise, structured outputs without needing specific guidance or prior examples.
The model's advanced capabilities were recently demonstrated through the creation of a highly detailed SVG graphic of a peacock. SVG, or Scalable Vector Graphics, is a code-based image format that requires the AI to precisely calculate coordinates and shapes rather than simply predicting pixels. This specific image was generated in a "zero-shot" fashion, meaning the model produced the complex graphic immediately without being given any examples to follow. To achieve this level of detail, the Gemini Pro Checkpoint utilized approximately 24,000 tokens and spent roughly six minutes in a state of high reasoning effort to ensure the accuracy and complexity of the final visual.
The public visibility of the model ID confirms that this checkpoint is active, signaling that Google is testing a seriously powerful new iteration of its technology. The ability to dedicate several minutes of processing time to a single, complex request marks a shift toward high-reasoning models that prioritize accuracy and depth over instantaneous responses. By successfully generating intricate vector art through pure reasoning, the Gemini Pro Checkpoint demonstrates a level of spatial and structural understanding that could eventually transform how AI assists in design, coding, and technical illustration.
07GPT image 2.5 Enhances Fidelity and Control
OpenAI has updated its image generation capabilities with the release of GPT image 2.5, making the process of creating and refining visuals significantly faster and more intuitive for the average user. The most immediate impact is a drastic reduction in waiting times, with the company claiming up to 50% lower latency, which refers to the delay between sending a request and receiving the result. This means users can move from a text prompt to a high-quality image much more quickly, streamlining the creative workflow and reducing the friction typically associated with AI-generated art.
Beyond speed, the update focuses on visual fidelity—the level of detail and realism in the output. ChatGPT image 2.5 produces sharper images with improved consistency and a more sophisticated handling of natural light and textures. One of the most practical improvements is the stronger subject preservation when using reference photos. This allows users to maintain the identity and characteristics of a specific person or object across different scenes, solving a common frustration where AI models often drift away from the original reference during the generation process.
The most significant shift in control comes from a new comment-based editing system. Previously, modifying a specific part of an AI image often required regenerating the entire piece or using complex tools, which frequently altered parts of the image the user wanted to keep. With this new system, users can modify exactly what they want by leaving comments on specific areas of the image. This enables more reliable, multi-round edits where only the targeted section changes while the rest of the composition remains untouched. By combining faster generation with this surgical level of control, OpenAI is moving the tool from a simple prompt-and-hope generator to a more precise editing suite that allows for iterative refinement.
08DeepSeek V4.1 Flash Launches with Open Weights
Users can now deploy a high-performance AI model on their own hardware without sharing sensitive data with a third party. DeepSeek has released V4.1 Flash as an open-weights model, meaning the underlying parameters that govern the AI's behavior are public and downloadable. This allows for private deployment on personal computers or mobile phones. To make this possible on consumer-grade hardware, the model supports quantization—a process that compresses the model to reduce its memory footprint. This is particularly valuable for users managing limited VRAM, the specialized high-speed memory required to run large models locally.
In terms of capability, V4.1 Flash is designed to compete with frontier models like GPT 5.6 Sol and Opus 5. It demonstrates a specific edge in the Cyber Gym benchmark, where it scored 84.5, outperforming GPT 5.6 Sol. However, performance varies by task; for instance, Opus 5 maintains a lead on terminal bench 4, scoring 51 compared to V4.1 Flash's 31. The model also features native visual understanding and an asymmetric architecture that utilizes 8 billion active parameters for input and 16 billion for output to increase overall efficiency.
One of the most distinct features of V4.1 Flash is its continuously controllable reasoning effort. While closed-source models from companies like GPT or Claude typically offer a few discrete settings—such as low, medium, or high—DeepSeek allows users to pass a specific numeric value from 1 to 100. This gives developers granular control over how much computational effort the model spends on a specific problem, allowing them to balance response speed and accuracy based on the complexity of the task.
This efficiency extends to hardware requirements, as the model significantly reduces the need for high-bandwidth memory and SSD storage. While its standing relative to GLM 5.3 remains uncertain, the combination of open research and accessible weights positions V4.1 Flash as a powerful tool for those seeking total control over their AI workflows and data privacy.
09Sensitive government and corporate data was allegedly leaked
High-stakes intelligence and corporate secrets may have been exposed through a hidden bridge between two major AI providers. Allegations have surfaced that DeepSeek has been proxying Anthropic's Claude model, effectively serving as a middleman that routes user requests to Anthropic while collecting the resulting exchanges to train its own systems. This arrangement has reportedly led to the leakage of highly sensitive information, including the full specifications and strategic objectives of a flagship AI program belonging to a technology company in the People's Republic of China. Even more critical is the reported exposure of data from a Russian government agency, specifically requests originating from an IT operator associated with the ministry of defense.
This proxying mechanism creates a dangerous security loophole where users believe they are interacting with one entity while their data is actually flowing to another. By routing requests to Claude, DeepSeek can offer users the high-performance capabilities of an "Opus level" model at a significantly lower price point. However, this cost-saving measure comes with a severe privacy penalty. The process allows for the harvesting of data, including chain-of-thought reasoning transcripts—the step-by-step logic a model uses to solve a problem—which can then be used to refine and distill model capabilities. This suggests that the ability to access premium AI performance for cheap often implies that the user's data is the actual currency being traded.
The risk extends beyond government agencies and large corporations to individual users. There are concerns that Anthropic retains data from personal accounts, which typically range in price from $20 to $200, despite public claims regarding privacy. While the company may anonymize this information to mask individual identities, the data is still maintained for internal use and model improvement. This pattern underscores a broader industry trend where the use of affordable or subsidized AI models often correlates with aggressive data retention policies. For users and organizations, the trade-off for lower costs is a loss of control over proprietary information and personal privacy.
10Apple announced two new Apple Watches featuring ambient list
Apple is redefining the role of the wearable device by integrating a more persistent form of artificial intelligence into its latest hardware. The company recently announced the Apple Watch Ultra 4 and the Series 12, both of which introduce ambient listening capabilities. This feature allows the devices to remain active and attentive to the user's surroundings at all times, potentially transforming the watch from a tool that responds to specific commands into a proactive assistant that understands the context of a user's day in real time.
The most critical technical aspect of this update is the commitment to local processing. While many AI-driven voice features rely on sending audio clips to distant servers for analysis, the ambient listening on the Apple Watch Ultra 4 and Series 12 runs entirely on the watch and the paired phone. By keeping the data processing local, Apple ensures that the audio captured by the device does not need to be sent to the cloud. This approach addresses a primary concern for users regarding privacy and data security, as the "always-on" nature of the microphone is mitigated by the fact that the information never leaves the user's personal hardware ecosystem.
This move signals a strategic bet on the form factor of the next generation of AI wearables. For some time, the industry has been uncertain whether the dominant AI interface would be smart glasses or a wrist-based device. By equipping both the premium Ultra 4 and the standard Series 12 with these capabilities, Apple is positioning the wrist as the ideal location for a constant AI companion. This shift moves the user experience away from the traditional "trigger-and-response" model toward a more seamless integration where the device can process environmental audio locally to provide immediate, relevant utility without the latency or privacy risks associated with cloud-based AI.
11Manual step confirmation in Astra can be automated using a verifying agent
Users often find themselves trapped in a tedious cycle of clicking confirmation buttons for every minor action an AI takes, which effectively defeats the purpose of automation. This friction can be eliminated in Astra by transferring the responsibility of step confirmation to a verifying agent. By adding a single line to the configuration settings, the system no longer requires the user to manually approve every individual progression. Instead, the verifying agent takes over the role of checking whether a step has been completed correctly, allowing the workflow to proceed without constant human intervention.
The necessity of this feature becomes clear when considering how AI models handle open-ended tasks. Without a mechanism to verify that a goal has been reached, Astra might continue to refine or "polish" a task indefinitely. This happens because the system lacks a point of reference to determine when the work is actually finished. By implementing a verifying agent, users provide the system with a way to cross-reference its progress against the desired outcome, preventing wasteful loops and reducing the amount of computational resources and tokens—the basic units of text processed by the AI—consumed during a session.
This shift toward autonomous verification complements Astra's broader ability to handle complex, repetitive data organization. For instance, the platform can be connected to Google Sheets through a one-click plug-in to automate the creation of structured directories. This directory can include specific categories, subcategories, and descriptions, transforming a manual research project into an automated process. By combining the power of verifying agents with these integration tools, Astra allows users to offload repetitive operational steps while retaining final decision-making power for critical permissions or financial expenditures.
12Codex Expands Context Window to 828,000 Tokens
Codex now allows users to significantly expand its context window—the amount of information the AI can process and remember at one time—from a default of 256,000 tokens to 828,000 tokens. This increase, which more than triples the default capacity, enables the model to handle far larger volumes of data or more complex instructions without losing the thread of the conversation. This expansion is achieved through a single line of configuration in a file located in the user's home folder.
This configuration can be handled automatically, as Astra can write the required key into the config file itself. It is important to note, however, that the actual usable token count is slightly lower than the configured 828,000. The system client must reserve a portion of the window for metadata and the model's own responses, meaning the full theoretical limit is not available for input data.
Despite this massive increase in general memory, Codex imposes strict limits on its "skill lists," which are the definitions of specific tools and capabilities the model can use. The system allocates only 2% of the context window or 8,000 characters to these skills. For a standard 250,000 token window, this means only about 5,000 tokens are available to cover all skill names, paths, and descriptions combined.
Overloading this skill list leads to a decline in performance. When too many skills are added, Codex truncates the descriptions to force them into the limited space. This leaves the model with only fragments of text, making it difficult for the AI to distinguish between different skills and choose the correct one for a task. While the overall context window has expanded, the rigid limit on skill definitions means that adding too many capabilities can actually make the model less effective at using them.
