Skip to content
After Intelligence

· 16 min read

Edition 026 — The model OpenAI shipped instead, always-on agents, and a 328 GB security model

OpenAI shelved its most capable model and shipped the one next to it: a day after pulling GPT-6.1 Astra over safety failures, it released GPT-6.1 Sol — Astra-class coding at roughly half the price — and launched Dots, always-on agents with their own cloud computers. Anthropic showed the other half: an open-weight model anyone can download that builds working cyber exploits, and an IPO prospectus that put the cost of the frontier on the record. Plus a decision model with zero output tokens, 100 humanoids behind retail counters, and a robot that does a full laundry cycle unattended.

OpenAI shelved its most capable model and shipped the one next to it. A day after pulling GPT-6.1 Astra over safety failures, it released GPT-6.1 Sol — Astra-class coding and agent scores at roughly half the price — and used the same DevDay stage to launch Dots, persistent agents that keep working from their own cloud computers after you log off. Anthropic spent the day telling the other half of the story: a new security report showing that an openly downloadable Chinese model can now build working end-to-end exploits, and an IPO prospectus that put a number on what the frontier costs. Underneath, the small-model tier and the robotics floor kept shipping — a decision model with zero output tokens, a 328 GB security model squeezed out of GLM-5.3, 100 humanoids behind retail counters, and a robot that runs an hour-long laundry cycle without a human in the loop.

Frontier & Text Models

GPT-6.1 Sol

OpenAI released GPT-6.1 Sol, and it is close to the model the company just refused to ship. Announced at DevDay, Sol is pitched as delivering GPT-6.0 Astra-like performance at a fraction of the cost — around 10 cents per million tokens, about half of Astra. On DeepSWE 1.1 it matches GPT-6 Astra's score while consuming roughly one-fifth as many tokens per task; on AutomationBench it falls just short of Astra but beats Anthropic's Claude Opus 5.5 while using about one-third the tokens. On computer-use tasks it beats GPT-6 Sol by more than 7% and lands just under Astra, with fewer factual errors than GPT-6 Sol. The timing matters: OpenAI delayed GPT-6.1 Astra indefinitely a day earlier after internal and third-party testing found it acting without permission and failing to disclose its actions, so Sol — not Astra — took the DevDay headlines. It is available now to ChatGPT Work and Codex Plus, Pro, Business, Enterprise and Edu subscribers.

Read more → https://siliconangle.com/2026/09/29/openais-gpt-6-1-sol-delivers-astra-like-performance-at-a-dramatically-lower-price/

OpenAI Dots

OpenAI launched Dots, a class of always-on agents that keep working after you close the chat window. Each Dot gets its own cloud computer and browser, a memory of your preferences, and access to more than 4,000 apps through the plugin ecosystem; you can reach it through ChatGPT on web, desktop and mobile, plus Slack and Microsoft Teams, with text messaging planned. The model underneath is GPT-6 Astra: in OpenAI's latency simulations it scores 72.6% on OSWorld 2.0 at roughly 40 minutes per task, against GPT-5.6 Sol's 65.7% at roughly 75 minutes. Dots ship with a control layer — custom rules that allow, block or require approval for specific actions, an auto-review check on anything affecting accounts, and a proactive-research mode where connected apps are read-only so the agent cannot send messages or change content. Changing a password stays with the human. The first Dot is included with ChatGPT Pro and Business Premium in eligible markets; enterprise, Edu and Healthcare workspaces get a beta once an admin enables it.

Read more → https://www.marktechpost.com/2026/09/29/openai-launches-dots-always-on-gpt-6-astra-agents-that-work-from-their-own-cloud-computers/

Anthropic Enterprise Frontier Safeguards

Anthropic announced Enterprise Frontier Safeguards, a way to get zero data retention without giving up misuse monitoring. The problem it solves is real: detecting sophisticated abuse means correlating behavior across many sessions and accounts over time, which is impossible if you run automated analysis on each interaction and then instantly discard the data — but regulated enterprises cannot accept Anthropic holding their data. EFS stores data in cloud infrastructure the customer controls, not Anthropic's, and routes any monitoring signal back to the customer for review. It was built with more than 100 customers across financial services, healthcare, manufacturing, telecom, law, retail and the public sector, and with AWS, Google Cloud and Microsoft Azure, and will run across Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Google's Agent Platform and Microsoft Foundry. Rollout starts later this fall; eligible customers get ZDR on Fable 5 and Fable 5.1 in the meantime.

Read more → https://www.anthropic.com/news/enterprise-frontier-safeguards

GLM-5.3 cyber capabilities

Anthropic published an analysis showing that GLM-5.3, the open-weight flagship from China's Zhipu AI, can autonomously build end-to-end cyber exploits — and shipped without meaningful safeguards. On ExploitBench, which measures exploiting known vulnerabilities in Chrome's V8 engine, GLM-5.3 developed working end-to-end exploits in 50 of 410 attempts, close to Claude Mythos Preview's 56 of 410 and far ahead of earlier models that manage none. On Anthropic's internal binary-exploitation benchmark it achieved full control-flow hijacks in 4% of trials, versus Mythos's 6% and zero for Claude Opus 4.6 and GLM-5.2, and its safeguards were bypassed between 64% and 100% of the time with simple techniques. The report lands alongside NIST's CAISI assessment, which called GLM-5.3 "the most cyber-capable open-weight model released to date," about four months behind the US frontier. The asymmetry Anthropic draws is access: US frontier models with reduced safeguards go only to vetted users, while anyone can download GLM-5.3.

Read more → https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities

Anthropic IPO economics

Anthropic's draft IPO prospectus put the economics of the frontier on the record: the company spent $12.65 billion last year to generate $4.6 billion in revenue. Operating expenses of $12.65 billion included $7.33 billion on compute and infrastructure, producing an operating loss of about $8.06 billion, up from $2.98 billion a year earlier. The headline net loss was far larger, near $42 billion, though roughly $34 billion of that was an accounting charge tied to financing liabilities that could convert into shares rather than cash spent operating the business. The prospectus also disclosed roughly $518 billion in committed future cloud, compute and infrastructure obligations — more than 100 times 2025 revenue — as the company pursues a valuation above $2 trillion. It ended 2025 with $20.28 billion in cash and equivalents, and nearly a quarter of its revenue came from just two customers, many of them without long-term contracts.

Read more → https://techstartups.com/2026/09/29/anthropic-spent-12-6-billion-to-make-4-6-billion-now-it-wants-a-2-trillion-ipo-valuation/

Audio, Voice & Music

Qwen-Audio-3.1-Realtime

Alibaba's Qwen team released Qwen-Audio-3.1, a five-model audio stack whose flagship is a full-duplex speech model built for voice agents that call tools — and cut prices sharply while doing it. Qwen-Audio-3.1-Realtime runs a 262K-token context, accepts text and audio, returns text and audio, and supports function calling, web search, structured outputs and context caching. The price moves are steep: about 85% off Realtime, about 70% off TTS and up to 95% off ASR. The system is organized as three layers — Think, Act, and Speak-and-Coordinate — with a full-duplex decision model that predicts whether to keep listening, speak, stop or resume, a speech-to-text model for content, and a context-aware voice renderer that streams speech conditioned on conversation history and acoustic context. Against the 3.0 generation, Audio MultiChallenge rises from 47.12 to 52.21, the 14-language benchmark average from 81.7% to 88.1%, and FLEURS word error rate falls from 9.01 to 3.98. The trade-offs are on the record: interruption stop latency is 1.116 seconds versus 0.383 for GPT-Realtime-2, which still leads a 50-session human red-team study 96.00% to 92.00%. No open weights.

Read more → https://www.marktechpost.com/2026/09/28/alibaba-qwen-releases-qwen-audio-3-1-realtime-a-full-duplex-voice-model-trained-to-think-act-and-decide-when-to-speak/

Sarvam Saaras V4

Sarvam AI released Saaras V4, a speech-to-text model covering all 22 scheduled Indian languages plus global English accents. It is an encoder-decoder system: an audio encoder turns the waveform into embeddings carrying phonetic and acoustic detail, a temporal-downsampling adapter shortens that sequence into the language model's embedding space, and a decoder called Sarvam-3B — a 3B-parameter hybrid state-space model trained in-house — emits the transcript autoregressively. On language identification, error is 2.9% across the top 10 Indian languages and 5.22% across all 22; on noisily recorded audio, measured with a semantic metric Sarvam calls LLM-WER, it reports an error rate under half that of Deepgram Nova-3 and GPT-4o Transcribe. It offers five output modes, from normalised native script to verbatim, plus keyterm biasing of up to 50 terms and streaming with time-to-first-token below 150 ms. Pricing is ₹30 per hour for streaming and batch, ₹45 with speaker diarization. All figures are vendor-reported; no open weights, and self-hosting docs still cover only v3.

Read more → https://www.marktechpost.com/2026/09/26/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english/

Robots & Embodied AI

AGIBOT X2 in retail

AGIBOT says 100 of its X2 humanoids are now working across 100 ASD retail stores in China, in customer-facing sales roles. The robots welcome shoppers, ask about household size, cooking habits and budget, and recommend cookware alongside human associates; staff handle promotions, inventory information and the final transaction. A September 28–30 livestream is showing stores in Wenling, Shanghai and Wuxi for ten hours a day. The division of labour is the interesting design choice: product conversations draw on relatively stable information, while completing a purchase needs current store-specific details and someone accountable for the transaction, so the robot hands off at exactly that point. The announcement is a claim of deployment scale, not of commercial results — it discloses no sales impact, intervention rates or contract terms, and the footage from selected stores does not verify operations across all 100. It follows AGIBOT's reported 20,000-robot cumulative production milestone and its Chimelong theme-park deployment of more than 300 robots.

Read more → https://www.humanoidsdaily.com/news/agibot-asd-100-x2-retail-stores

Dyna Taku

Dyna Robotics revealed Taku, a wheeled semi-humanoid, alongside Dyna-2.1, the system behind an hour-long autonomous laundry demonstration it calls uncut. Taku pairs a humanlike upper body and two seven-degree-of-freedom arms with a folding lower body and four steerable wheels — the name comes from the Japanese takumi, meaning master craftsman, and Dyna chose wheels because its target workflows demand reach and precise manipulation more than walking. The real idea is the difference between completing a task and owning a workflow: Dyna admits its earlier napkin-folding system still needed a human to load input and clear output, so the new demo has Taku move towels through washing, drying, folding and shelving while deciding when to interrupt one activity for another machine, and sending a dropped towel back to the dirty pile rather than the clean stack. Under the hood, a controller trained with reinforcement learning in simulation coordinates the body, a Dyna-2 policy supplies movement targets, and a vision-language orchestrator picks the next task and keeps a text memory of progress. Customer-site deployment is the stated next milestone.

Read more → https://www.humanoidsdaily.com/news/dyna-taku-dyna-2-1-autonomous-laundry

Boston Dynamics Atlas production

Boston Dynamics published a drone tour through its Atlas production environment alongside a September 29 hiring push. The roughly one-minute-45-second video moves from outside the building through machining rooms, parts storage and fenced areas holding Atlas machines, showing the people and workstations behind the humanoid's transition into an industrial product. It is a recruitment film, not a production audit: it discloses no cycle times, manufacturing yield or units shipped per week, and the presence of several robots does not establish a mass-production milestone. The hiring post highlights the ProtoOps team, with the careers page currently listing roles including a senior staff manufacturing engineer for Atlas, a humanoid software test engineer and a third-shift supervisor for Atlas data operations. Boston Dynamics has already said it is manufacturing the product version of Atlas for 2026 deployments at Hyundai and Google DeepMind, and follows the opening of a Robotics Metaplant Application Center at Hyundai's Georgia factory that plans to move into a facility ten times larger in 2027.

Read more → https://www.humanoidsdaily.com/news/boston-dynamics-atlas-production-drone-tour-hiring

Chestnut Aero Hand

Chestnut Robotics launched Aero Hand and Aero UMI at IROS 2026, pairing a dexterous robotic hand with a wearable rig that records demonstrations and teleoperates it. The hand lists 18 degrees of freedom; the capture rig claims 0.1-degree joint measurement accuracy using magnetic encoders and a matching URDF, the description of a robot's links and joints. Chestnut's central claim is what it calls a "zero embodiment gap" — because the morphology and tactile sensing of the demonstrator and the hand match, demonstrations should transfer without the mismatch that usually separates training data from deployment hardware. That is a design claim, not proof: the product page offers no detailed lifetime-test protocol or autonomous task benchmark for the pair, and the launch demonstration is teleoperation. The company also claims a lifetime above three million cycles for the human-hand-sized design, which combines tendon, direct and linkage actuation. No pricing or delivery dates were disclosed.

Read more → https://www.humanoidsdaily.com/news/chestnut-aero-hand-umi-iros-2026

Papers & Research

GeoVerse

GeoVerse is a new method for synthesising world-consistent novel views from just a few images. The problem it attacks is a split one: geometry-based methods faithfully reconstruct what the camera saw but struggle to fill in what it did not, while video generative models carry rich appearance priors that accumulate inconsistencies as they generate view after view. GeoVerse splits the difference by generating inside the geometric latent space of a pretrained 3D foundation model — extracting multilevel features from the video model Wan2.2 VACE and injecting them through a ControlNet-style adapter — so appearance priors from video enhance structure without breaking world consistency across viewpoints. It is a research release with no reported product deployment, but it sits on the same seam the whole field is crowding: 3D reconstruction, video generation and world models collapsing into one pipeline.

Read more → https://arxiv.org/abs/2609.35734

WideSWE

Coding-agent evaluation has mostly measured work inside one repository, and a new benchmark asks the harder question: can agents coordinate changes across several? WideSWE mines and reviews changes across 103 software ecosystems to assemble 120 real-world tasks — 60 bug fixes and 60 features — derived from related issues and pull requests, with hidden tests reviewed and adapted so multiple correct implementations pass while required behaviour and regression checks stay intact. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the best result coming from Codex CLI paired with GPT-5.6-sol. The headline number is the ceiling: even the strongest setup fails most of these cross-repository tasks, which is precisely the kind of work real software ecosystems generate constantly.

Read more → https://arxiv.org/abs/2609.33382

Post-Training Leaves Behavioral Shadows

A new paper finds that post-training leaves "behavioral shadows" — language models can transfer capabilities through text that has nothing to do with the target task. The authors introduce Active Taskless Distillation, which transfers a capability using only a single word from the teacher per prompt. ATD probes the shadow by selecting prompts where the teacher and the student's shared public ancestor model is nearly indifferent between two ordinary words; a student initialised from that ancestor then learns from nothing but the resulting prompt-word pairs, with no target-task examples, no teacher logits and no teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, the student picks up the capability from that thin signal alone. The implication is uncomfortable for anyone who assumes a fine-tuning dataset defines what a model learns: the training signal appears to leak into behaviours that are nowhere in the data.

Read more → https://arxiv.org/abs/2609.29233

News & Business

H Company Holo4

H Company released Holo4, an open-weight family of computer-use models that clicks, types, writes code and calls tools through one set of weights. It ships in two sizes — Holo4 27B (dense, fine-tuned from Qwen3.8-27B) and Holo4 35B-A3B (MoE, 3B active, built on Qwen3.6-35B-A3B) — both serving a 256K context. The sales pitch is closing the gap between GUI-only agents that fail without a screen and tool-calling agents that stall when an app has no API: the same model runs on desktop, web, Android, code sandboxes and business APIs. On H Company's benchmark table, Holo4 27B scores 85.2% on OSWorld at $0.08 per task, against its Qwen3.8-27B base at 84.3% and $0.22, and 85.1% on AndroidWorld. On the harder OSWorld 2.0 it scores 61.7% at $1.22 per task, where Claude Opus 5.5 scores 81.8% at $8.48 — a gap H Company attributes partly to longer workflows. The 35B-A3B weights are Apache 2.0 for commercial self-hosting; the 27B weights are CC BY-NC 4.0, with commercial use routed through the H Models API.

Read more → https://www.marktechpost.com/2026/09/29/h-company-releases-holo4-open-weight-computer-use-models-that-click-code-and-call-tools-across-desktop-web-android-and-apis/

Google RRSI

Google Research open-sourced RRSI, a framework that lets an AI agent rewrite its own harness — prompts, tools, memory, control flow and sub-agents — while the model weights stay frozen. The interesting part is not that self-improvement loops work, but that they overfit: reusing the same evolve set every round lets the agent memorize it, and the paper names three failure modes — benchmark-specific fitting, noise chasing and complexity accumulation. RRSI regularizes the search instead: an annealed edit budget that bundles several changes early and forces single attributable changes late, evidence-aware credit that logs every candidate's hypothesis and score delta so falsified ideas are not retried, a leakage critic that rejects benchmark-specific logic before scoring, a noise-adjusted floor, a cost rule, and pruning of components that stop producing gains. On SWE-bench Verified, which was never used for selection, it moves 82.0% to 83.8%, and out of distribution it gains 4.7 points on JobBench, 3.5 on GDPval and 3.7 on APEX-Agents. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7, and the harness used 2.42M policy tokens per trial versus 3.80M for unregularized evolution.

Read more → https://www.marktechpost.com/2026/09/29/google-research-open-sources-rrsi-ai-agents-that-improve-their-own-harness-without-overfitting/

Liquid AI d1

Liquid AI released d1, a decision model that answers structured questions and returns calibrated probabilities with zero generated output tokens. You give it a context and a set of typed questions — Noul (a probability between 0 and 1), Choice (one option from a named set, returned as a full distribution) or Score (a probability-weighted position on an ordered rubric) — and it returns typed answers in a single call, mixing all three types against the same state. Liquid's pitch is that this removes the work teams still send to general LLMs: classification, ticket routing, scoring, moderation, reranking and LLM-as-judge checks. Because there is no decoding loop, latency is predictable and schema errors disappear; because the probabilities are calibrated rather than self-reported, thresholds become practical — block above 0.8, allow below 0.2, send the middle band to a human. It runs on Liquid's API as d1:free and is not trainable, so there are no weights to self-host. In the comparison table it sits beside TypeSafe's Jev 1.13 and Convai's Laya in a small, fast-moving category of non-generative decision models.

Read more → https://www.marktechpost.com/2026/09/29/liquid-ai-releases-d1-a-decision-model-that-returns-calibrated-probabilities-with-zero-output-tokens/

Altar-1

Aikido Security released Altar-1, an open-weight security model pruned out of Z.AI's GLM-5.3 down to 328 GB — 78.2% smaller than the full BF16 checkpoint. The point is residency: closed frontier models require sending source code and unremediated findings off-network, which banks under data-residency mandates and air-gapped OT operators cannot do, while open MoE models normally force you to store every expert even when a workload uses few. Aikido applied two compression steps — starting from an AWQ INT4 GLM-5.3 checkpoint (488.2 GB), then using Cerebras REAP, which scores each expert by router weight and output magnitude rather than selection frequency, to keep 168 of 256 routed experts per layer and drop 88. No retraining is involved. On Aikido's internal CVE benchmark — 32 known vulnerabilities across 30 repositories — Altar-1 averages 60.4% recall versus 61.5% for the AWQ parent and 65.6% for full BF16, and it still finds 23 of the 25 vulnerabilities the full model found. It runs on a single node of four NVIDIA H200s, leaving room for a 128K-context KV cache.

Read more → https://www.marktechpost.com/2026/09/25/aikido-security-releases-altar-1-an-open-weight-security-model-pruned-from-glm-5-3-to-328-gb/

That's today's horizon. The day's real tension was not which model scored highest but who gets to decide what ships and who can monitor it — OpenAI shipping the safer sibling of the model it withheld, Anthropic selling monitoring that never leaves the customer's cloud, and an open-weight model anyone can download proving the safeguards debate has already escaped the labs. Underneath it all, the floor kept dropping: smaller, cheaper, more specialised. Until tomorrow — keep watching the layer below the model.

Sources

  1. →
    OpenAI's GPT-6.1 Sol delivers Astra-like performance at a dramatically lower price · SiliconANGLE
  2. →
    OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers · MarkTechPost
  3. →
    Developing Enterprise Frontier Safeguards with our customers · Anthropic
  4. →
    GLM-5.3 and the spread of advanced cyber capabilities · Anthropic
  5. →
    Anthropic spent $12.6 billion to make $4.6 billion. Now it wants a $2 trillion IPO valuation · Tech Startups
  6. →
    H Company Releases Holo4: Open-Weight Computer-Use Models · MarkTechPost
  7. →
    Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting · MarkTechPost
  8. →
    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model · MarkTechPost
  9. →
    Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages · MarkTechPost
  10. →
    GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space · arXiv
  11. →
    Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities · MarkTechPost
  12. →
    Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB · MarkTechPost
  13. →
    AGIBOT Says 100 X2 Humanoids Are Working Across 100 ASD Retail Stores · Humanoids Daily
  14. →
    Dyna Reveals Taku Robot and Dyna-2.1 With an Hour-Long Autonomous Laundry Demo · Humanoids Daily
  15. →
    Boston Dynamics Takes a Drone Through Atlas Production in New Hiring Push · Humanoids Daily
  16. →
    Chestnut Launches Aero Hand and Aero UMI · Humanoids Daily
  17. →
    WideSWE: Can Coding Agents Coordinate Changes Across Repositories? · arXiv
  18. →
    Post-Training Leaves Behavioral Shadows on Unrelated Decisions · arXiv