Skip to content
After Intelligence

· 11 min read

Edition 028 — The frontier's best model ships to defenders first, and a $42 billion loop around Anthropic

Google handed Gemini 4 Argon to cyber defenders before the public, Broadcom agreed to lend Anthropic up to $42 billion to lease its own chips, and Micron posted a record quarter as AI memory demand reshaped the chip market.

The frontier's newest systems arrived behind guarded doors this week: Google handed its most capable model to cyber defenders before the public, and Broadcom agreed to lend Anthropic up to $42 billion to lease the very chips it helps build. Around them the industrial layer thickened — memory posted a record quarter, Meta's agent app outran ChatGPT's early download pace, and small, fast decision models kept eating the chatbot's job.

Frontier & Text Models

Google's Gemini 4 Argon release

Google unveiled Gemini 4 Argon, its most capable model yet — and gave it first to cyber defenders. Argon supports output of up to 1 million tokens, up from 64,000 in earlier Gemini models, letting a single run continue through far longer reasoning and execution. Google says it scored 77.9% on DeepSWE v1.1 for long-horizon software engineering, a new high, and 68% on CWE-bench v1, tying for first place on vulnerability remediation. It also ranked first on the Vals Index across economically significant work and posted 91.7% on LVBench for long-video understanding. Access starts through the Fairwind Program for trusted defenders, with introductory API pricing of $2 per million input tokens and $10 per million output tokens before rising to $4 and $20. Google says teams used Argon agents on data-center telemetry to free more than 300 tebibytes of memory.

Read more → https://techstartups.com/2026/10/01/google-launches-gemini-4-argon-its-most-advanced-ai-model-yet/

Broadcom and Anthropic compute financing

Broadcom agreed to lend Anthropic up to $42 billion to lease its own chips. Disclosed in Anthropic's IPO prospectus and first reported by Reuters, the financing could cover roughly one-third of the AI startup's $125.2 billion five-year commitment to lease tensor-processing-unit capacity. Broadcom works with Google on custom TPUs and is expected to become Anthropic's largest compute customer in 2027. The arrangement is one of the clearest examples yet of the circular financing spreading across the industry — a supplier financing its own customer's demand. "It feels that there's quite a concentrated bet right now on two companies being able to generate enough revenues to support all the financing that's happened," Rothschild & Co's Robert Leitao told Reuters.

Read more → https://techstartups.com/2026/10/01/broadcom-to-lend-anthropic-up-to-42-billion-to-lease-its-own-chips-as-ai-circular-financing-hits-a-new-level/

Meta Muse AI app

Meta's Muse AI assistant passed 5 million U.S. downloads just 22 days after launch, outpacing ChatGPT, Grok and Claude. Sensor Tower estimates Muse crossed the milestone faster than ChatGPT (56 days), Grok (103 days) and Claude (492 days). It remains at the top of the U.S. Apple App Store charts. Unlike the chatbot Meta spent years trying to make competitive, Muse is built as a personal agent that can browse websites, send emails, fill out forms, book travel and work across connected services, running inside a dedicated virtual machine in Meta's cloud and continuing after the user hands it a task. Sensor Tower notes a giant asterisk: by late September, Muse was receiving as much as half of Meta's daily house-ad impressions.

Read more → https://techstartups.com/2026/10/01/metas-muse-ai-hits-5-million-downloads-in-22-days-outpacing-chatgpt-grok-and-claude/

OpenAI said it shut down a coordinated distillation campaign it tied to associates of Moonshot AI. The company described the effort as adversarial distillation aimed at extracting protected chain-of-thought reasoning. It began at low volume in early July, then spiked to 16,000 requests from more than 4,000 users on July 24–25, part of a broader cluster of over 15,000 accounts that OpenAI says it fully disrupted by July 28. The episode is a concrete case of a frontier lab treating reasoning traces as a defended asset rather than a public byproduct.

Read more → https://techstartups.com/2026/10/01/top-tech-news-today-october-1-2026-amd-broadcom-google-lg-meta-micron-openai-reddit-more/

Video & World Models

PixelUMM paper

PixelUMM is an encoder-free model that unifies image and video understanding and generation directly in pixel space. Unified multimodal models usually rely on separate visual representations for understanding and generation, inflating visual context length; pixel-space modeling offers an encoder-free alternative but had not extended cleanly to video. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with modality-specific paths — one visual interface for both reading and making video.

Read more → https://arxiv.org/abs/2609.38597

Audio, Voice & Music

Qwen-Audio-3.1-Realtime voice model

Alibaba's Qwen team released Qwen-Audio-3.1, a five-model audio stack whose centerpiece is a full-duplex speech model built to call tools. Qwen-Audio-3.1-Realtime runs two models behind one voice — a decision model that predicts whether to keep listening, speak, stop or resume, and a speech-to-text model — so it can hold a conversation without waiting for turn-taking gaps. Context is 262K tokens, with 245K max input and 16K max output, and a single request can carry a 1M-token-class budget per minute under default limits. It is live as a managed API on QwenCloud over WebSocket, with no open weights. Qwen also cut prices hard: about 85% on Realtime, about 70% on TTS, and up to 95% on ASR.

Read more → https://www.marktechpost.com/2026/09/28/alibaba-qwen-releases-qwen-audio-3-1-realtime-a-full-duplex-voice-model-trained-to-think-act-and-decide-when-to-speak/

Tacit-TTS paper

Tacit-TTS removes the transcript requirement from zero-shot voice cloning — and makes it over 10x faster. Autoregressive semantic TTS systems clone voices well but decode slowly; non-autoregressive alternatives are faster but usually need a transcript of the reference speech at inference. Tacit-TTS, distilled from IndexTTS2, replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, adds training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets it matches competitive zero-shot quality while generating speech more than 10x faster than IndexTTS2.

Read more → https://arxiv.org/abs/2609.38658

Robots & Embodied AI

Boston Dynamics Atlas hands

Boston Dynamics gave Atlas four-finger hands, raising each hand from seven to 13 degrees of freedom. The new GR3 hand adds a more capable opposable thumb with four degrees of freedom, while each of the other three fingers has three; the fingers can spread apart and the thumb moves along and across them to support different pinches and tool grips. The design is meant to move the humanoid beyond grasping objects toward manipulating them in-hand and operating human-scale tools. Boston Dynamics deliberately omitted a fifth finger to limit cost, size and failure points, and built the hardware for accurate simulation and reinforcement learning.

Read more → https://www.humanoidsdaily.com/news/boston-dynamics-atlas-four-finger-hands-13-dof

Agile Robots Hitachi manufacturing

Agile Robots and Hitachi are partnering to commercialize physical AI on the factory floor. The collaboration pairs Munich-based Agile Robots' Agile ONE humanoid and AgileCore platform — simulation, control and orchestration — with Hitachi's AI and industrial-integration expertise, including Hitachi's low-power edge AI semiconductors for processing visual and tactile information. Target tasks include parts handling, machine loading, assembly and production changeovers, the operations where workpiece variation makes fixed automation hard. Hitachi will route commercialization through HMAX Industry. The announcements disclose no deal value, robot order quantities or deployment timetable.

Read more → https://www.humanoidsdaily.com/news/agile-robots-hitachi-physical-ai-manufacturing

Agility FORT Digit 5 safety

Agility Robotics and FORT Robotics are extending Digit 5's safety systems beyond the robot itself. A new memorandum of understanding, signed October 1, covers an Offboard Safety Bridge designed to connect the humanoid to external safety systems, complementing Digit 5's planned onboard human detection, safety cues and safe motion control. The goal is to let safety inputs come from both the robot and its surroundings as a mobile humanoid moves between work areas. FORT already supplies technology used in earlier Digits, and Agility is also working with NVIDIA on the onboard safety architecture. No delivery date or new certification was announced.

Read more → https://www.humanoidsdaily.com/news/agility-fort-digit-5-offboard-safety-bridge

Small & On-Device Models

MiniCPM5-2B on-device model

MiniCPM5-2B packs long context and tool-calling into a 2-billion-parameter on-device model. OpenBMB's release targets edge deployment with conversational use in English and Chinese, and it has drawn about 1,700 likes and more than 940,000 downloads since shipping in early September. It is trained on OpenBMB's Ultra-FineWeb and UltraData corpora, including a dedicated agent-SFT mixture. It is the kind of small, capable model that keeps making the case that "on-device" no longer means "toy."

Read more → https://huggingface.co/openbmb/MiniCPM5-2B

Papers & Research

LOCI world-model paper

LOCI gives video world models a spatial memory that does not blow up with video length. A world model should reproduce a region when the camera revisits it, which needs both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches keep detail but grow with length; recurrent memory is compact but compresses history into a fixed state. LOCI is a hybrid: in half the transformer blocks attention keeps a key-value cache, while in the other half it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry — so viewpoint enters both memory addressing and stored content.

Read more → https://arxiv.org/abs/2609.40222

Looped MoE scaling laws paper

A new paper writes the first scaling law that models recurrence and sparsity together. Looped transformers raise computational depth at fixed parameters, while mixture-of-experts sparsity expands total capacity at fixed active compute — two complementary routes to efficient scaling that prior scaling laws treated in isolation. The work introduces Loop Scaling Laws, built around a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises it. The laws predict held-out loss of looped models more accurately than prior alternatives and recover the standard dense and MoE scaling laws as special cases, giving designers a principled footing for looped MoE models under compute and memory constraints.

Read more → https://arxiv.org/abs/2609.40316

Agent Error Dataset paper

The Agent Error Dataset turns 50,000 failed agent rollouts into training signal. An unsuccessful agent rollout holds more information than its final reward — the observations, the actions chosen and the environment's responses. AED comprises 50,228 error-diagnosis pairs drawn from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models in text-based agent systems, retaining source traces and execution metadata so failures can be re-diagnosed without repeating the rollout. Its five-stage Agentic Error-to-Training pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence — a large, reusable corpus for error-aware post-training.

Read more → https://arxiv.org/abs/2609.40111

News & Business

Micron posted a record fiscal fourth-quarter of $54.23 billion in revenue as AI memory demand reshaped the chip market. That is up from $41.46 billion in the prior quarter and $11.32 billion a year earlier. GAAP net income reached $37.70 billion, and operating cash flow climbed to $43.97 billion. For the full fiscal year, Micron generated $133.19 billion in revenue. The numbers underline how far memory has moved from a commodity to a bottleneck that AI infrastructure spending now pulls hard against.

Read more → https://techstartups.com/2026/10/01/top-tech-news-today-october-1-2026-amd-broadcom-google-lg-meta-micron-openai-reddit-more/

Cloudflare Clef decision models

Cloudflare released Clef and Clef-flash, the first models trained by its Workers AI team — decision models that return typed probabilities instead of text. Each reads an input state and a schema of typed questions and returns a probability for every allowed answer, with no free-form text and nothing to parse. Clef supports three question types: yes/no, single-choice with per-option probabilities, and rubric scoring. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI's Jev API; a single Workers AI request can carry up to 64 questions and four images. Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, and both keep the backbone's vision encoder.

Read more → https://www.marktechpost.com/2026/10/01/cloudflare-releases-clef-and-clef-flash/

Cohere Embed 5

Cohere released Embed 5, an embedding family for enterprise search, RAG and agentic retrieval. It ships in two tiers — Embed 5 Pro for maximum retrieval quality and Embed 5 Fast for latency and cost on the query path — both accepting text, images and fused text-plus-image inputs, covering 100+ languages and reading up to 128K tokens. The key design choice is that Pro and Fast share one embedding space, so you can index with one and query with the other; normalized to Pro-plus-Pro at 100, a Pro index queried with Fast scored 98.4. Outputs run 256 to 2048 dimensions in float, int8 or binary. Pro costs $0.12 per million text tokens, Fast costs $0.08, and image inputs cost $0.40 per million tokens on both.

Read more → https://www.marktechpost.com/2026/10/01/cohere-releases-embed-5/

Kev-9B decision model

Kev-9B is an open decision model that answers typed questions with calibrated probabilities and zero output tokens. It is a LoRA adapter on Qwen3.5-9B-Base, released under Apache 2.0, and built on the same "System One" pattern as TypeSafe AI's Jev. On a locked, read-once out-of-domain test of six never-trained public sources and held-out policy structures (656 questions) it scored 85.2% accuracy with a Brier score of 0.199; it hit 83.4% on a hard held-out set and 79.1% on developer-tooling questions. Because it emits a probability for each allowed answer rather than a sentence, the result needs no parsing — a small model doing the classification work a chatbot does badly.

Read more → https://huggingface.co/jaredpalmer/kev-9b

That's this week's horizon. Reply with the model, paper or story you want us to take into the next edition.

Sources

  1. →
    Google launches Gemini 4 Argon, its most advanced AI model yet · Tech Startups
  2. →
    Broadcom to lend Anthropic up to $42 billion to lease its own chips · Tech Startups
  3. →
    Meta's Muse AI hits 5 million downloads in 22 days · Tech Startups
  4. →
    Top Tech News Today, October 1, 2026 · Tech Startups
  5. →
    Cloudflare Releases Clef and Clef-flash · MarkTechPost
  6. →
    Cohere Releases Embed 5 · MarkTechPost
  7. →
    Kev-9B: open decision model card · Hugging Face
  8. →
    MiniCPM5-2B: on-device model card · Hugging Face
  9. →
    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime · MarkTechPost
  10. →
    Tacit-TTS: Efficient Transcript-Free Voice Cloning · arXiv
  11. →
    LOCI: Spatial Linear Memory for Streaming World Models · arXiv
  12. →
    PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation · arXiv
  13. →
    Scaling Laws for Looped Mixture of Experts · arXiv
  14. →
    Agent Error Dataset: Scaling 50,000 Error-Diagnosis Pairs · arXiv
  15. →
    Boston Dynamics Gives Atlas Four-Finger Hands · Humanoids Daily
  16. →
    Agile Robots and Hitachi Partner on Physical AI for Manufacturing · Humanoids Daily
  17. →
    Agility and FORT Extend Digit 5's Safety Systems · Humanoids Daily