Skip to content
After Intelligence

· 9 min read

Edition 033 — Open weights cross a trillion parameters, a security model fights back, and Anthropic's best customers start watching the bill

Mistral ships a 1.05-trillion-parameter open-weight flagship, Reflection AI's Beam bets on inference efficiency, an open security model beats Claude Opus on cost by 31x, and Anthropic's two biggest customers quietly trim their Claude use ahead of a $2-trillion IPO — plus Suno Speech, Microsoft's streaming STT, EmbeddingGemma 2, and humanoids from Agility and Minerva.

The open-weight frontier crossed a trillion parameters today — and it came with a price tag, a security mission, and a set of nervous customers. Mistral shipped a 1.05-trillion-parameter flagship while Reflection AI put a 501B efficiency play on the waitlist, Cantina released an open model that hunts bugs and beats Claude Opus on cost by 31x, and Anthropic's two biggest corporate customers quietly cut their Claude use weeks before a $2-trillion IPO. The infrastructure underneath kept growing too — Google locked in 3.59 gigawatts of power — even as AI-assisted attackers hit seven South Korean banks. Here is the edition.

Frontier & Text Models

Mistral Large 4 Le Chonk

Mistral's "Le Chonk" goes public preview with a 1.05T-param granular MoE. The model activates 49B/token (~4.7% of weights) and pairs a 1.6B vision encoder with a 1M-token context, trained on 3,800 Grace Blackwell GPUs in Mistral's own EU datacenters across >160 languages with native image input. API pricing is $1.36/1M input and $4.18/1M output; weights ship end of October with no self-hosting yet. Cyber results: 93% Cybench, 82% CyberGym-E2E, while several closed frontier models score near zero because they refuse the task.

Read more → https://www.marktechpost.com/2026/10/06/mistral-ai-releases-mistral-large-4-le-chonk-a-1-05t-parameter-open-weight-multimodal-moe/

Reflection AI Beam

Reflection AI ships "Beam," its first open-weight model, targeting inference efficiency over raw scale. The sparse MoE has 501B total / 23B active params for coding, reasoning and agents, trained on 23.8T pretrain tokens after curation removed ~95% of raw internet tokens but kept ~1.8T high-quality tokens filters would drop. Reflection claims it competes with larger open models like GLM 5.2 at 3-4x less inference compute on reasoning benchmarks. It is not yet self-hostable (final red-teaming, waitlist); Kimi K3 stays ahead on raw capability.

Read more → https://www.marktechpost.com/2026/10/05/reflection-ai-introduces-beam-a-501b-open-weight-moe-model-with-23b-active-parameters-for-coding-and-agentic-workloads/

DeepSeek funding

DeepSeek nears a $12B round, blowing past its ~50B yuan target. >80B yuan ($12B) is committed with Tencent and CATL among investors; Bloomberg reports the final tally could approach 100B yuan ($15B). The round completes in October, with a possible Shanghai STAR IPO early 2027 (CITIC Securities hired). DeepSeek hit $1B annualized revenue <2 weeks ago — >2x a few months earlier — and sought capital in July at ~$74B valuation after a prior ~$7.4B round.

Read more → https://techstartups.com/2026/10/06/deepseek-to-raise-12-billion-in-tencent-and-catl-backed-funding-ahead-of-ipo/

Anthropic customers

Anthropic's biggest customers are pulling back ahead of its IPO. Meta's Claude Code users fell to ~30,000 from ~60,000 earlier in 2026 (The Information). Microsoft cut projected internal Anthropic spend by >1/3; its >60,000-person cloud/AI group dropped per-employee allowances from up to $100K/month to ~$10K/month in most cases, and a ≥$1B/yr self-use projection fell >1/3. A draft prospectus showed two unnamed customers equal to 24% of 2025 revenue; the IPO could value Anthropic >$2T.

Read more → https://techstartups.com/2026/10/06/meta-cuts-claude-users-by-50-as-microsoft-slashes-ai-budgets-from-100k-to-10k-per-employee-ahead-of-anthropic-ipo/

Cantina apex-flash-1

Cantina and Yeta Labs release apex-flash-1, an open-weights vuln-research model under MIT license. It is an RL fine-tune of Z.ai GLM-5.3-Flash with 321.3B total params (18B active base), trained on 150 tasks from 50 real vulnerability cases — 72% auth/identity/scope flaws, 18% accounting/numerical. On 60 held-out tasks it solved 40/60 (66.7% pass@1) for ~$2.38, versus base GLM-5.3-Flash at 36/60 ($4.56) and Claude Opus 5 High at 43/60 (71.7%) for $74.68 — ~31x the cost. That works out to ~$0.06 per solved task vs $1.74 for Opus.

Read more → https://www.marktechpost.com/2026/10/04/can-an-open-model-do-security-research-cantinas-apex-flash-1-solves-40-of-60-held-out-bug-tasks/

Video & World Models

JEPA-Anything

JEPA-Anything applies one world-model recipe across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather. The work comes from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton, and extends JEPAs with Orthogonal Predictive Factorization. It splits the latent target of width d into K subspaces of width r (d=K×r, usually K=4), each with a dedicated predictor, recombined via the Moore-Penrose pseudoinverse. This fixes the capacity-allocation problem where high-variance structure dominates and weaker modes receive conflicting gradients.

Read more → https://www.marktechpost.com/2026/10/05/beyond-domain-specific-world-models-jepa-anything-uses-1-recipe-for-7-fields/

Audio, Voice & Music

Suno Speech beta

Suno opened its "Speech" beta to everyone on web and mobile on 1 Oct 2026 after roughly a month of closed testing, free during beta. It generates a spoken voice and its background music together as one unified track rather than classic TTS: type an idea or script, then describe the voice and the music. This is distinct from Suno's earlier "Voices" feature (Aug 2026), which worked from recorded vocals — Speech generates the speech itself from text, no recording needed.

Read more → https://jackrighteous.com/blogs/guides-using-suno-ai-music-creation/suno-speech-beta-spoken-audio-background-music

MAI-Transcribe-2-Streaming

Microsoft AI launched MAI-Transcribe-2-Streaming, its first streaming STT model, on 1 Oct 2026 alongside TTS models MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for both final and first-partial transcript accuracy across 60 languages with automatic continuous language detection; first partial hypotheses arrive just over 100ms after audio. AA measured final transcript at 2.5% WER at 0.13s after end of speech (#1) and first partial at 2.5% WER at 0.12s (#1). Runners-up: Grok Voice Transcribe 2.0 (2.7%/0.49s) and Muse Voice Transcribe (3.1%/0.16s).

Read more → https://www.marktechpost.com/2026/10/02/microsoft-ai-releases-mai-transcribe-2-streaming-1-real-time-speech-to-text-model-on-artificial-analysis/

Robots & Embodied AI

Reka Rho-1

Reka's Rho-1 is a 19B omni-reasoning model trained from scratch that unifies text, images, video and robot actions in one network. Text, vision and actions all become tokens in a single context window, replacing pipelines of modality-specific models. In one unedited 5-turn session it drew a lighthouse, boxed it, animated it, edited the video into a snowstorm, and explained the difference — with no tool call and no second model. It is currently a research preview.

Read more → https://www.marktechpost.com/2026/10/05/reka-releases-rho-1-a-19b-omni-reasoning-model-that-understands-generates-video-and-outputs-robot-actions-in-one/

Agility Digit 5

At its 6 Oct Investor Day, Agility detailed the economics behind its $2.5B SPAC merger with Churchill Capital Corp XI (Nasdaq AGLT, expected to close this quarter). Digit 5 costs ~$150,000 in parts at launch, above Digit 4's ~$125,000, and the $300M/1,000-robot order unlocks in tranches, with each proven use case releasing a batch over 3 years of service across 4. GA was described as end of 2027 "and early 2028," with early-access units shipping H1 2027 (the S-4 had pointed to late 2026). RaaS robots sit outside ~$8M annual capex, and the CFO wants sub-10% debt once public.

Read more → https://www.humanoidsdaily.com/news/agility-investor-day-digit-5-launch-costs-and-a-300m-order-that-unlocks-skill-by-skill

Minerva Roger

Minerva raised ~$10M in pre-seed on 6 Oct to build "Roger," a humanoid letting specialists do dangerous work from a safe distance. General Catalyst led, with Long Journey Ventures and Credo Ventures co-leading; targets are oil/gas and public safety, including ordnance disposal and hazardous-material response. Roger uses shared autonomy: a remote human operator supplies high-level intelligence while onboard AI handles balance, fall recovery and navigation. A first paid oil-and-gas pilot runs this fall, operational deployments are targeted for 2027, and the current robot is a prototype.

Read more → https://www.humanoidsdaily.com/news/minerva-roger-humanoid-10-million-hazardous-work

SuperCLUE Embodied Brain

GPT-6 Astra leads overall among ten models in SuperCLUE's September 2026 EmbodiedCLUE-VLA "Embodied Brain" evaluation (reported 6 Oct), scoring 86.75 overall. Its clearest edge is interaction & planning: 85.71 vs runner-up Gemini-3.8-Flash's 64.29, a 21.42-point gap (4.63 overall). Qwen3.8-Max-0902 and Nebula-EmbodiedBrain tie at 77.48, while international models appear as references outside numbered rankings. Scores assess cognitive capability, not demonstrated robot task success.

Read more → https://www.humanoidsdaily.com/news/superclue-september-2026-robot-brain-planning-gap

NVIDIA Cosmos3-DROID

A 5 Oct developer walkthrough details building a streaming robotics learning pipeline with NVIDIA Cosmos3-DROID. It turns robot data collection and world-model training into a continuous stream rather than batch jobs, aimed at robot policies trained on NVIDIA's Cosmos world model.

Read more → https://www.marktechpost.com/2026/10/05/building-a-streaming-robotics-learning-pipeline-using-nvidia-cosmos3-droid/

Small & On-Device Models

EmbeddingGemma 2

Google DeepMind released EmbeddingGemma 2 on 6 Oct, a 740M-param open multimodal embedding model built on Gemma 4 that embeds text, code, images, video and audio into one 768-dim space with 8K-token context under Apache 2.0. It targets on-device search, classification and privacy-first RAG, and is modular: a 270M text/code backbone (130M transformer + 140M embedder), a 170M optional vision encoder, and a 300M optional audio encoder — load 270M/440M/570M/740M, all sharing one vector space. Weights are on Hugging Face and Kaggle, with Ollama, llama.cpp GGUF and LiteRT builds available now.

Read more → https://www.marktechpost.com/2026/10/06/google-deepmind-releases-embeddinggemma-2-a-740m-open-multimodal-embedding-model-built-on-gemma-4/

Papers & Research

HLA-WM

A 5 Oct arXiv paper (2610.05739) proposes HLA-WM, applying hybrid linear attention to long-horizon video world models. It targets the compute/memory bottleneck that limits world models as they roll out over long horizons.

Read more → https://arxiv.org/abs/2610.05739

News & Business

Investigators see signs of AI-assisted attacks on seven South Korean financial firms. >67,000 people are believed exposed across Shinhan Bank, KB Kookmin Bank, Hana Bank, BNK Busan Bank, Yegaram Savings Bank, Welcome Savings Bank and Hyundai Capital. Traces of ARTEX AI, a Chinese-language open-source autonomous pentest tool, were found on infrastructure believed used. Authorities have not identified the attacker, and the evidence does not establish the tool developer's or the Chinese government's involvement.

Read more → https://techstartups.com/2026/10/06/top-tech-news-today-october-6-2026-amazon-amd-apple-deepseek-google-uber-more/

Google signs a 3.59-GW power deal with Constellation Energy, one of the largest US corporate power deals. 890 MW comes from upgrades at 11 existing nuclear units in Illinois, Pennsylvania and New Jersey under a 20-year PPA, with first added nuclear output expected in 2028. Another 2,700 MW is long-term within PJM, not tied to a specific source. Constellation says the deal could support >$4.3B in investment.

Read more → https://techstartups.com/2026/10/06/top-tech-news-today-october-6-2026-amazon-amd-apple-deepseek-google-uber-more/

Laya

Convai Innovations' Laya, an open-source decision engine released 6 Oct, was one of the most-starred ML repos of September 2026. It is a non-autoregressive "System 1" model: a 421M-param encoder reads text plus typed questions (a choice between labels, a score on a scale, or a yes/no) and returns a probability for every option in one forward pass with zero output tokens. Positioned as the open answer to TypeSafe's Jev, its pitch is speed and calibrated probabilities. A developer guide exercised it on the CLINC150 banking intent dataset, covering zero-shot accuracy, option-order sensitivity, calibration, temperature fitting, an abstention gate, and pydantic-schema outputs.

Read more → https://www.marktechpost.com/2026/10/06/a-developers-guide-to-laya-zero-shot-decisions-and-calibration/

That's this week's horizon. Until tomorrow — keep watching the parameters, the power bills, and the postmortems.

Sources

  1. →
    Mistral AI releases Mistral Large 4 'Le Chonk' · MarkTechPost
  2. →
    Reflection AI introduces Beam, a 501B open-weight MoE · MarkTechPost
  3. →
    DeepSeek to raise $12 billion ahead of IPO · Tech Startups
  4. →
    Meta cuts Claude users by 50% as Microsoft slashes AI budgets · Tech Startups
  5. →
    Can an open model do security research? Cantina apex-flash-1 · MarkTechPost
  6. →
    Top Tech News Today, October 6, 2026 · Tech Startups
  7. →
    Reka releases Rho-1, a 19B omni-reasoning model · MarkTechPost
  8. →
    JEPA-Anything: 1 recipe for 7 fields · MarkTechPost
  9. →
    Suno Speech beta: spoken audio with background music · Suno Guides
  10. →
    Microsoft AI releases MAI-Transcribe-2-Streaming · MarkTechPost
  11. →
    Google DeepMind releases EmbeddingGemma 2 · MarkTechPost
  12. →
    A developer's guide to Laya: zero-shot decisions and calibration · MarkTechPost
  13. →
    Agility Investor Day: Digit 5 launch costs and a $300M order · Humanoids Daily
  14. →
    Minerva: Roger humanoid, $10 million · Humanoids Daily
  15. →
    SuperCLUE September 2026 robot-brain planning gap · Humanoids Daily
  16. →
    Building a streaming robotics learning pipeline with NVIDIA Cosmos3-DROID · MarkTechPost
  17. →
    HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models · arXiv