The frontier spent the day proving two things at once: that open weights are catching up fast, and that speed cuts both ways. Three open models landed inside a week — Mistral's 1.05-trillion-parameter "Le Chonk," Reflection's 501B Beam, and Reka's 19B Rho-1 — while security researchers showed that a single chat message to one public agent could hand over every AI agent in an AWS account. Underneath, the money kept moving: Microsoft and Meta quietly pulled back from Claude, Manus raised half a billion dollars after Beijing blocked its sale, and a UK institute warned that open models now trail closed ones in cyber capability by only four months. Here is the edition.
Frontier & Text Models

Mistral released Mistral Large 4, a 1.05-trillion-parameter multimodal MoE it nicknamed "Le Chonk," trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters. Only about 49B parameters — roughly 4.7% of the total — activate per token, which is how a trillion-class model serves at mid-tier pricing: $1.36 per million input tokens and $4.18 per million output. It carries a 1.6B-parameter vision encoder, a 1M-token context window, and support for more than 160 languages including every official EU language. Mistral's standout numbers are in cybersecurity, where it reports 93% on Cybench and 82% on CyberGym-E2E, and notes that several closed frontier models score near zero because they simply refuse the task. The API is live now as a public preview, but the weights do not ship until the end of October.

Reflection AI introduced Beam, its first open-weight model — a 501B-parameter sparse MoE with 23B active per token, aimed squarely at coding and agentic workloads. Beam was pretrained on 23.8 trillion tokens, and Reflection says its curation threw out about 95% of raw internet text while keeping roughly 1.8 trillion high-quality tokens that ordinary filters would have discarded. The architecture interleaves local and global attention with fine-grained routed experts, building on DeepSeek-V3's auxiliary-loss-free load balancing; the busiest expert ended pretraining at just 1.04x average load. Reflection is candid that Kimi K3 stays ahead on raw capability, so Beam's pitch is efficiency — it claims to compete with larger open models like GLM 5.2 at 3 to 4 times less inference compute on reasoning benchmarks, with a user-facing reasoning-effort dial. It is not self-hostable yet: Beam is in final red-teaming and early access runs through a waitlist.
Some mathematicians are calling for a boycott of OpenAI after it flooded their field with hundreds of machine-generated proofs. After OpenAI published hundreds of AI-generated mathematical manuscripts, the Association for Human Mathematics accused the company of violating core norms of scientific research and urged mathematicians to stop working with it. The group's chair, Fields Medalist Terence Tao, shared the statement as a guest post on his blog while outlining a "Math 2.0" era. The dispute turns on a real question: whether a problem can count as solved before its proof has been independently worked through and made understandable to humans. Many of the papers are so dense and oddly written that even leading experts say they need AI help to decipher them, if they can at all.

Anthropic added two beta features to Claude that push it from chat answers into live artifacts: Dashboards and Motion. Dashboards connects data sources like BigQuery, Databricks, Snowflake or Salesforce and turns them into auto-updating live dashboards from a text prompt, with every number showing its underlying query and results pipeable into tools like Amplitude, Grafana or Hex. Motion generates animated explainer videos from text, diagrams and images, editable as code and exportable as MP4. Dashboards is available to paid users and Motion to Team and Enterprise plans — a quiet signal that the productivity battle is shifting from "can it write" to "can it ship a working thing on your data."
![]()
A new memory layer gives a language model a 50-million-token working window that is faster and cheaper than recomputing the prompt. The public package, galahad-kv, saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back byte-exact without recomputation; run on 50 million tokens of real public text through vLLM on a single H100 with Gemma 4 12B and 31B, every one of 100 probed blocks loaded back with no recompute at depths from 0 to 50M tokens. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, with GPU memory flat across the whole stream. Asked about facts planted millions of tokens earlier, the 12B model answered correctly 82 times out of 100 and the 31B model 98 out of 100 — and neither made an answer up. The limits are explicit: this reuses stored state rather than widening attention, one block loads at a time, and the store needs terabytes of NVMe.
Read more → https://arxiv.org/abs/2610.10845
Image & Vision

Google released Nano Banana 2.1 for image generation and editing, and cut prices roughly in half against the model it replaces. A 1K image now costs 3.36 cents, down from 6.70 cents under Nano Banana 2, with Google saying it improves on previous versions "across the board." The model runs on Gemini 3.6 Flash rather than the newer Pro stack, so Google's top image quality still sits with Nano Banana Pro — and the next big jump is expected with the announced Gemini 4 Argon. The half-price cut is the story: image generation is rapidly becoming a commodity input for high-volume developer pipelines rather than a premium feature.
Read more → https://the-decoder.com/googles-new-image-model-nano-banana-2-1-generates-better-images-for-less-money/

Black Forest Labs released Flux 3 Image, the image side of its Flux 3 family, built around multi-step edits that leave the rest of a picture untouched. It covers text-to-image, image-to-image, text rendering and photorealism, lets users compose scenes with bounding boxes and up to ten reference images, and outputs up to 4K. API access was half price through October 8, and companies can license commercial weights to run and fine-tune the model on their own infrastructure. An open-weight version is expected in the coming weeks — the same pattern BFL has followed across the Flux line, where the paid API leads and the weights follow.
Audio, Voice & Music

Suno, best known as an AI music generator, launched a feature called Speech that produces spoken text and matching background music as a single audio track. Users type an idea or written text and describe the voice and music style they want, and the model generates both the voice and the sound together — pitched for poems, meditations and bedtime stories. Product chief Jack Brody says the company tested it with a small group for a month, and the beta still has rough edges; a British accent can sometimes come out Australian. The launch lands as AI music generators face mounting copyright pressure — major labels have sued Suno, and a Munich court recently ruled against the startup, rejecting fair use for copyrighted training data.
Read more → https://the-decoder.com/ai-music-generator-suno-can-now-create-spoken-audio-with-matching-background-music/

Microsoft AI released a new real-time transcription model and two text-to-speech models aimed squarely at voice agents. MAI-Transcribe-2-Streaming covers 60 languages, delivers its first partial results in just over 100 milliseconds — enough for an agent to respond while someone is still mid-sentence — and Microsoft says it ranks first for accuracy on Artificial Analysis; an hour of audio costs $0.54 at the introductory price through year-end. MAI-Voice-2.1 speaks 23 languages in the same voice with a native accent in each, while the Flash variant hits 150-millisecond latency at $15 per million characters instead of $22. Both voice models can clone a voice from a few seconds of reference audio, with built-in safeguards meant to prevent misuse.
Read more → https://the-decoder.com/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice-agents/
Robots & Embodied AI

Reka released a research preview of Rho-1, a 19B omni-reasoning model trained from scratch that collapses text, images, video and robot actions into a single network. Most multimodal systems are pipelines — a planner hands jobs to specialists, adding latency and narrow context at every handoff; Rho-1 removes those handoffs by treating text, vision and robotic actions as tokens inside one context window. Every input and output uses one of two native formats: discrete tokens for language, symbolic reasoning and high-level commands, and continuous tokens for image latents, video frames, robot actions and proprioception. Each transformer block holds two expert weight streams that share attention over the same KV cache, with a discrete handoff token letting the generation stream render pixels from the accumulated state. In an unedited session the model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm and explains the difference — all in five turns, with no tool call and no second model.

Nvidia is pitching its Halos full-stack safety system for physical AI, covering compute, sensors, software and inspection tools for machines that operate near people — and it now reaches humanoids and robotaxis. Agility Robotics has already built Halos into its Digit humanoid so that safety sensing lives on the robot rather than in external work-cell hardware. Nvidia previously extended the architecture from autonomous vehicles to warehouse robots, humanoids and surgical systems, and CEO Jensen Huang has described physical AI as a multi-billion-dollar revenue line. The strategic logic is straightforward: a standardized safety layer lowers the certification burden for robotics startups that cannot build in-house compliance stacks, while keeping Nvidia inside the robotics bill of materials beyond GPUs alone.
Papers & Research
![]()
Tetris3D reconstructs a full 3D scene from a single image by conditioning each object's generation on the geometry and physical relationships of its neighbors. Existing methods tend to generate objects independently or couple them only implicitly, which gives little guidance about whether neighboring objects that interact with one another actually fit. Tetris3D instead forces each object's shape and pose to stay geometrically and physically plausible within the scene, and the authors back it with ComOb, a physics-simulation dataset of 1.2 million scenes with per-object meshes and pairwise physical-relation annotations. On synthetic and real scenes it recovers coherent shapes and poses even when interacting regions are occluded, and posts state-of-the-art results in both generation quality and physical stability.
Read more → https://arxiv.org/abs/2610.10539
![]()
A new paper argues that text-to-music systems should be judged against a per-request rubric, not a single opaque relevance score — and builds an agent to optimize for it. A global text-audio score can overlook the implicit intent in underspecified prompts and mask failures in instrumentation, structure, rhythm or mood progression. The authors reformulate text-to-music alignment as satisfying independently verifiable rubric items covering both explicit requirements and implied musical intent, and instantiate it as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. Their MIRA agent then grounds a request into rubrics and searches over prompt revisions for a black-box generator under a bounded budget, generating music, verifying it against the rubrics and using the feedback to guide a trajectory-aware tree search.
Read more → https://arxiv.org/abs/2610.10355
![]()
Long-WAM shows that giving a world-action model more visual history only helps if its video backbone was pretrained autoregressively. The framework scales the context of causal world-action models under real-time control constraints, and its central finding is that access to history is not the same as using it. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain autoregressive pretraining pushes peak success higher still. Long-WAM also leads compared methods on LIBERO-Long, RoboTwin 2.0 and DOMINO, and runs on RTX 5090, DGX Spark and Jetson AGX Thor at 107.4 ms per action chunk including future-video latent prediction — with real-time deployment on the Unitree G1 and YAM, including 95% success on dynamic cup stacking.
Read more → https://arxiv.org/abs/2610.10528
![]()
RoboJEPA is a robotic world model that finally gives the field a way to predict how latent world models scale. Built on the Joint Embedding Predictive Architecture and trained on a dataset spanning 12 robotic embodiments, it shows that imagination error — the error of its latent rollouts — follows a second-order power law in compute, which lets the authors predict model quality well beyond the scale where they fit the law. Downstream planning performance improves predictably with compute and correlates strongly with imagination error, making it a usable proxy for real-robot evaluation instead of an expensive hardware test. The authors also deploy the latent world model zero-shot as a robotic agent that plans toward a single goal image to solve long-horizon tasks on real hardware, and release all checkpoints and deployment code.
Read more → https://arxiv.org/abs/2610.10515
News & Business

JetBrains released Mellum2.1, a 12B mixture-of-experts model that activates just 2.5B parameters per token and ships under Apache 2.0. The architecture is unchanged from Mellum2 — 28 layers, 64 experts with 8 active — and the upgrade comes almost entirely from reinforcement learning in real software environments, producing a small, self-hostable model that explores a repository, edits files and checks its own changes. It beats Mellum2 on 15 of 17 listed benchmarks and wins 5 of 17 against Qwen3.5-9B, with its best result an 82.0 on LiveCodeBench v6 (ahead of Qwen3.5-9B's 75.4) and its worst a 17.4 on Terminal-Bench 2.1. It runs in 131,072-token context on your own GPUs via vLLM or SGLang, with GGUF builds starting at 7.0 GB for llama.cpp, Ollama and LM Studio.
Read more → https://www.marktechpost.com/2026/10/08/jetbrains-releases-mellum2-1-a-12b-moe-open-model-for-coding-agents/

Perplexity released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models that retrieve across text, images and rendered PDF pages in one shared space. The two sizes split neatly by role: a 0.6B model (about 340M active parameters for images) meant as a live query encoder that runs fully local on a laptop or small GPU, and a 9B model for building the document index. Because they share one embedding space, a 9B index can be searched with 0.6B queries, recovering about half the quality gap at a fraction of the cost; a single query reaches 92.4% on MADQA. Its 128-dimension token vectors are 16x to 32x narrower than rivals at 2,048 to 4,096 dimensions, and both models are MIT-licensed on Hugging Face. The catches: it stores one vector per token so indexes grow with document length, a single input cannot mix text and images, and all scores are self-reported with the technical report still pending.
A single chat message to one public agent was enough to hijack every AI agent in an AWS account, Zenity researchers found. A chain of vulnerabilities in Amazon's Bedrock AgentCore platform let the researchers take over all of a company's AI agents in the same AWS region: the agents lacked proper isolation and handed over internal AWS credentials when asked, and the platform's broad default permissions let them reach source code, passwords, private conversations and the long-term memory of other agents. AWS has partially fixed the issue by making it harder for new agents to retrieve internal metadata and by tightening the default execution role, but the researchers still recommend that companies manually assign permissions rather than trust defaults. It is a compact illustration of the year's recurring lesson: the security boundary of an agent fleet is only as strong as its weakest public-facing agent.
That's this week's horizon. Open weights keep marching toward the frontier while the money, the money fights and the failure modes move faster still — a trillion parameters on one side, one prompt on the other. Until tomorrow.