The frontier got cheaper this week, not smarter — and that is the story. OpenAI slashed GPT-6 Sol and Luna API prices by up to 58% and Anthropic shipped Claude Opus 5.5 at 40% lower running cost than Opus 5, two moves that put near-flagship capability at commodity prices. Underneath them, the open-weight tier kept closing: Xiaomi put a 1.02T-parameter omnimodal model on public weights, robotics moved from demonstration videos into production lines and rehearsal-based control, and a Brazilian lab trained a usable decision model for about $104. Rounding it out: speech models that out-reason their base, a 100M-parameter diarizer that tracks eight overlapping voices, and Google's plan to fly four TPUs in orbit on October 1.
Frontier & Text Models
GPT-6 gets two cheaper siblings as OpenAI slashes API prices by up to 58%. OpenAI released GPT-6 Sol and GPT-6 Luna, two models that sit below the flagship Astra but are trained with the same methods. Sol targets complex coding and professional work; Luna targets fast, high-volume jobs. Their API prices were cut by 50% on input against GPT-5.6 promotional pricing (58% on Luna's output) — Sol to $2 input and $10 output per million tokens (from $4 and $20), Luna to $0.10 input and $0.50 output (from $0.20 and $1.20, a 58% cut on output). On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2% at $0.27 per task, against Claude Opus 5's 26.9% at 11.1x that cost; on DeepSWE v1.1 Sol at max effort reaches 68.8%, 1.1 points behind Claude Fable 5 at about 80% lower cost per task.

Anthropic's Claude Opus 5.5 matches its bigger sibling at 40% lower running cost. Opus 5.5 is the first model in Anthropic's Claude 5.5 family, and the company says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 at default settings. On Anthropic's own benchmarks it leads on agentic coding, computer use and knowledge work: Terminal-Bench 4.0 climbs to 66.4% (Fable 5.1 scored 55.8%, Opus 5 52.3%, GPT-6 Astra 57.9%), FrontierCode v1.1 to 54.4%, and OSWorld 2.0 to 81.8%. It is a managed API model only — available on the Claude Platform, AWS, Google Cloud and Azure, with no self-hosting — and Anthropic cautions that benchmark margins are becoming a less reliable guide than cost-adjusted scores.
Read more → https://www.marktechpost.com/2026/09/22/anthropic-claude-opus-5-5-release/

Kyutai's Voice of Reason does maths out loud, without a text model in the loop. Voice of Reason is a pair of open-weight speech-to-speech models that answer maths questions as audio, built by fine-tuning GLM-4-Voice-9B with supervised fine-tuning and reinforcement learning. There is no transcription step and no separate text LLM. On spoken GSM8K, accuracy rises from 27.3% for the base model to 77.1%; SFT alone — trained on 150,616 Orca-Math problems rewritten for speech and voiced in many voices — already reached 61.7%. The team calls it the first application of RL to maths reasoning in speech-native models. Both BF16 checkpoints ran on a single H100.

Julia 1 is a 144M decision model trained for about $104. Supersonic Labs, a small Brazilian lab, released Julia 1 — a compact decision model rather than a chatbot. You hand it context, a question and between 2 and 20 candidate answers; it picks one and returns a probability for every option. It handles three decision types through one API: choice (pick a label), score (expected index on an ordered rubric) and noul (the probability a yes/no statement is true). It starts from JHU CLSP's mmBERT-small, a 140M-parameter multilingual encoder trained on over 1,800 languages, keeps the encoder and tokenizer, and adds a decision head. Total cloud GPU spend for training and experiments was about R$540 (US$104.08); the FP32 weights occupy 550.5 MiB, and an ONNX build runs in the browser via WebGPU.

Qwen-Image-2.1 puts generation and editing in one 7B model that fits on a desktop card. Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image and image-editing model whose visual generation component holds just 7B parameters across 32 single-stream DiT layers. It generates regular or transparent RGBA images from text, edits transparent layers and extracts subjects from photographs, accepts up to 10 reference images, and supports local edits specified by circles, painted annotations or separate masks while preserving the identity of people and products. A combination of mixed-granularity attention and prefix KV cache reuse keeps compute low enough to run locally — Tom's Hardware reports it producing images on hardware like an RTX 3090 in under 30 seconds — and benchmark coverage places it as an open-weight contender against OpenAI's and Meta's image models.
Read more → https://huggingface.co/Qwen/Qwen-Image-2.1

Audio, Voice & Music
Nemotron 3 Diarization tracks eight overlapping speakers in real time with 100M parameters. NVIDIA's open-weight speaker-diarization model answers the one question ASR leaves open: who spoke when. It doubles the four-speaker ceiling of NVIDIA's earlier streaming Sortformer checkpoint, handles voices that overlap, and uses a single checkpoint for both offline recordings and live streaming. It runs in 16 kHz mono, converts audio to Mel-spectrograms on a 10 ms step, and feeds a 31-layer Transformer encoder. The weights ship under the OpenMDW License 1.1, which permits commercial use, and run on Linux through NVIDIA NeMo on Ampere, Ada Lovelace, Hopper or Blackwell GPUs.
Read more → https://www.marktechpost.com/2026/09/23/nvidia-releases-nemotron-3-diarization/

Robots & Embodied AI
World Action Agent lets a vision-language model rehearse robot actions before taking them. WAA is a multi-agent harness that has general-purpose VLMs pilot robots using basic tools, with every decision made inside a visual action workspace. The workspace does three things: it selects contact views automatically from scene geometry, so the model sees the scene around the current interaction; it turns each action into an editable proposal that the agent — or an Imagination Agent — previews and revises against planning feedback; and it closes the loop with in-view correction, letting the agent cancel residual offsets in the view where it observes them. The same workspace also lets WAA pick up embodied procedural knowledge from expert videos and human teaching.
Read more → https://arxiv.org/abs/2609.29964
![]()
AGIBOT's 20,000th robot rolls off the line, doubling its cumulative output since March. The Shanghai company announced the milestone on September 24, adding 5,000 units since its June production update and presenting the step from 15,000 to 20,000 as evidence of a shift from demonstrating robot capabilities to deploying them at scale. Its graphic breaks the total into 5,638 A-Series, 8,757 X-Series and 5,605 G-Series machines. The 20,000th unit, identified in Chinese reporting as an A3 Ultra, went to the Chimelong theme park, where AGIBOT has parked more than 300 robots. The count is company-reported cumulative production and does not by itself establish how many machines have been sold or put into regular service.
Read more → https://www.humanoidsdaily.com/news/agibot-20000-robots-production-milestone

Tesla is reportedly building hundreds of Optimus robots a week — but not the ones it will sell. The Information reports Tesla increased Optimus production roughly tenfold over recent months, from a few dozen a week in the second quarter to several hundred a week in August, according to Humanoids Daily's read of the report. Most of those machines are used internally for testing, training and data collection. The V3 units now being manufactured differ from the eventual customer version, which still has to clear tougher durability and reliability requirements, and the report describes hand-assembly difficulties, sensor failures and inconsistent supplier quality. Tesla reportedly holds more than 500,000 hours of training data and plans to lease robots to selected commercial customers as part of a fleet-learning strategy.
Read more → https://www.humanoidsdaily.com/news/tesla-optimus-production-hundreds-per-week-report

Asimov open-sources the training code behind its humanoid's walk. The Menlo Research project released Asimov 1's locomotion policy and the code used to train it on September 25, extending its earlier hardware and simulation releases. The public repository contains an Isaac Lab training framework, reward definitions, actuator settings and simulated variations intended to help controllers cope with real hardware, and the team describes the package as a starting point for adapting the controller to hardware changes and exploring different walking styles. One caveat worth stating plainly: the announcement names a trained policy checkpoint, but Humanoids Daily could not locate a downloadable checkpoint in the linked training repository or its GitHub releases at the time of review.
Read more → https://www.humanoidsdaily.com/news/asimov-open-source-locomotion-policy-training

Papers & Research
ViRDM generates streaming video in a few steps without a teacher or a critic. Few-step autoregressive video diffusion normally leans on Distribution Matching Distillation, which needs a large pretrained teacher plus an online critic to estimate distributional error. ViRDM asks whether that stack can be dropped entirely by post-training only the generator against a precomputed target distribution, transferring representation distribution matching from one-step image generation to causal video. The paper names three barriers it had to clear — a memory-intractable gradient path, a video-specific optimisation regime, and representation distributions that underconstrain motion — and addresses them with stochastically truncated clean-exit supervision, a lightweight VAE decoder and staged vector-Jacobian products.
Read more → https://arxiv.org/abs/2609.28923
![]()
A cognitive-science exam finds object permanence is still missing from most video models. Researchers built WROP, a data infrastructure of 150 hand-designed, cognitive-science-inspired tasks split into six categories, with Blender generators that randomise speed, lighting and camera angle while preserving each task's structure — yielding 10,000+ samples per task, a 1.5M-sample training corpus and a 300-question exam. They then evaluated 14 video models on that exam: three reference-to-video, seven edit and four continuation models, alongside their own 16B world model PWM-WROP. The framing is the point: video generators are treated as a paradigmatic class of world model, so measuring whether they hold an object in mind while it is out of view is a test of physical reasoning, not image quality.
Read more → https://arxiv.org/abs/2609.28654
![]()
News & Business
Xiaomi's MiMo-V2.6-Pro-RL puts a trillion-parameter omnimodal model on open weights. MiMo-V2.6-Pro-RL is a sparse mixture-of-experts model with 1.02T total and 42B activated parameters, a 1M-token context, and native text, image, video and audio input. Its selling point is training method: Xiaomi scaled reinforcement-learning compute with a single mixed run across coding, general agents, visual tasks and cybersecurity — "You Only RL Once" — rather than separate per-domain runs. Groupwise agentic grading (rubric synthesis plus advantage redistribution) then ranks passing trajectories and steers toward shorter, cheaper solution paths. On the model card's own numbers it posts 71.9 on DeepSWE v1.1, 53.1 on AutomationBench v1.0.6 (ahead of Claude Opus 5's 50.3) and 94.0 on CyberGym.
Read more → https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
![]()
Saaras V4 covers all 22 scheduled Indian languages in a single speech-to-text model. Sarvam AI's new recogniser is an encoder-decoder system: an audio encoder turns the waveform into embeddings, a downsampling adapter shortens the sequence, and a 3B-parameter hybrid state-space decoder — trained from scratch in-house — emits the transcript. It adds global English accents to its Indic coverage, and Sarvam reports the lowest average word error rate among the models it benchmarked across seven English datasets (six drawn from Hugging Face's Open ASR Leaderboard). On Vistaar it reports results for 10 Indian languages using both WER and LLM-WER, the latter separating real meaning errors from harmless Indic-script spelling variants. Deployable through Sarvam's API today; the weights are not public.

GLiNER2.5-Decide is a 340M decision model that answers structured questions on a CPU. Fastino Labs' new open-weight model takes text plus a schema of typed questions and returns structured answers, each carrying a probability distribution, a confidence score and constraint-feasibility metadata. It targets the constant judgment calls inside agent pipelines: routing, triage, tool selection and guardrails. It is deliberately not generative — a DeBERTa-v3-large encoder fine-tuned from gliner2-large-v1 produces no tokens and needs no prompt template; the encoder scores every permitted answer for the text and schema, then a constrained decoder searches for the highest-scoring joint assignment the declared rules allow. The weights ship under Apache 2.0 and install with pip install gliner2, running on CPU, GPU or air-gapped machines.

Exa's Agent Ultra sends a swarm of subagents to build lists that other deep-research systems cannot finish. Agent Ultra is the highest effort level of Exa's Agent API, built for research that has to run to exhaustion: large list building, entity enrichment, and questions needing thousands of sources. It splits a task into subtasks, assigns subagents to research several domains at once, and routes frontier models only to the steps that need them. Exa reports that Ultra beats Opus 5.5, GPT-6 Astra and Perplexity Agent — each at maximum effort — across four research benchmarks: WANDR soft recall 81.4% versus 72.3%, 26.0% and 40.1%; DeepSearchQA F1 93.9%; WideSearch row-level F1 58.9%; and 2,451 passing entities per task on Company Find-All against 146, 113 and 98. Complex tasks typically finish in about 30 minutes, with the hardest taking up to three hours.

Altar-1 squeezes a 753B security model down to 328 GB so it can run inside the firewall. Aikido Security's first open-weight model is a compressed version of Z.AI's GLM-5.3, a 753B mixture-of-experts model that routes each token to 8 of 256 experts per layer, or about 40B active parameters. Two compression steps get it down: AWQ INT4 quantisation, which stores routed expert weights in 4 bits with 16-bit activations while attention, the shared expert, dense layers and the head stay in BF16; then Cerebras REAP expert pruning, which scores each expert by router weight and output magnitude rather than only routing frequency. The result runs with vLLM on a single node of four NVIDIA H200s and powers Aikido Machine, the company's autonomous pentesting appliance for on-prem and air-gapped networks — the point being that security context, source code and unremediated findings never leave the network.

Google plans to put AI hardware in orbit on October 1. Project Suncatcher, Google's effort to test whether AI compute can operate in space, will launch its first experimental satellite on October 1 — a refrigerator-sized spacecraft carrying four Google Tensor Processing Units and running Gemini workloads hundreds of miles above Earth. Developed with Planet and flying aboard SpaceX's Transporter-18 rideshare mission, the test will collect real-world data on how the TPUs respond to launch stresses, radiation and extreme temperature swings. The mission is tiny compared with a terrestrial data centre, where thousands of accelerators work together, but Google frames the larger question as whether AI compute could eventually move beyond electrical grids, land constraints and water-intensive infrastructure on Earth.

That's today's horizon. If one thread runs through all eighteen items, it is cost: half-price frontier APIs, a decision model trained for $104, a trillion-parameter reasoning model on open weights, and a security model small enough to sit inside the firewall.