Skip to content
After Intelligence

· 15 min read

Edition 017 — Type-Safe Brains, Fence-Free Robots, and Video That Thinks in Physics

Two frontier systems abandon chat entirely, a 3B voice model arrives on under 70,000 hours of speech, interactive video finally holds 720p, a humanoid is cleared to work beside people without a fence, and a 7B model keeps pace with 235B classmates.

Today's issue runs the full stack: two new frontier systems that abandon chat entirely, a 3B voice model built on under 70,000 hours of speech, real-time interactive video that finally holds 720p, a humanoid cleared to work beside people without a fence, and a 7B model that keeps pace with 235B classmates. The through-line is efficiency — fewer tokens, fewer hours, fewer watts, and less of the web memorized.

Frontier & Text Models

Gemini 3.8 Live Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, native speech-to-speech models for real-time voice agents, live in the Gemini Live API and Google AI Studio. These are hosted models with no open weights, positioned as a streamlined alternative to the cascaded ASR-plus-LLM-plus-TTS pipelines that most voice products still assemble by hand. Gemini 3.8 Live targets scale and cost, while Extended Thinking adds multi-step reasoning while speaking. That reasoning shows up on the scoreboard: Extended Thinking takes first on Artificial Analysis' Speech-to-Speech Quality Index at 82.6, leads agentic task completion with 68.6% on tau-Voice and 35.1% on Sierra's tau-Voice-banking, and hits 97.7% on Big Bench Audio, while Gemini 3.8 Live took second in the Speech Agent Arena human-preference evaluation. The Live API exposes asynchronous function calling — tool calls run in the background while audio keeps streaming — plus live visual context and alphanumeric precision across 97 languages. Read more → https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/

GPT-6 Astra Drone Bench Andon Labs put OpenAI's GPT-6 Astra through Vending-Bench 2 and Drone-Bench, and the results expose a widening gap between frontier models on long-horizon autonomy. In Vending-Bench 2 each model receives $500 and must run a simulated vending machine business for a year; across six runs Astra averaged a $15,515 final bank balance against Claude Fable 5.1's $5,422 — and Fable's best run at $9,874 still fell short of Astra's worst at $13,272. Astra is the first OpenAI model to top the Vending-Bench 2 leaderboard, and the gap to second place is the largest the benchmark has recorded. On Drone-Bench, Astra is the first model whose best attempts beat the human-AI baseline on all five subtasks, including writing code that lets a drone autonomously find and follow a specific person — though its success rate remains unreliable. Read more → https://the-decoder.com/gpt-6-astra-pilots-a-surveillance-drone-and-runs-a-business-on-its-own/

Audio, Voice & Music

Cohere North Small Translate Cohere released North Small Translate, a 218B open-weight translation model that returns to the exact task the Transformer was introduced on in 2017. Built with Cohere Labs and language-tech partner RWS, whose Language Weaver scientists shaped quality, it is a decoder-only sparse Mixture-of-Experts with 218B total and 25B active parameters, 128 experts with 8 activated per token plus shared experts, and a sigmoid router over expert logits. It covers 50 languages from Albanian to Vietnamese and scores 83.6 averaged across all of them on Cohere's WMT26 evaluation — a figure Cohere says beats DeepL and Google Translate as well as open options like GLM 5.2 and Mistral Large 3. It is the first translation model in Cohere's North family, following Tiny Aya and Command A Translate, and it is free on Cohere's API until rate limits, self-hostable non-commercially, with a commercial license available. Read more → https://www.marktechpost.com/2026/09/10/cohere-releases-north-small-translate-a-218b-moe-translation-model-that-scores-83-6-on-wmt26-across-50-languages/

rumik-oss 1 capabilities rumik ai released rumik-oss 1, a 3B multilingual text-to-speech model trained on fewer than 70,000 hours of speech that performs competitively with far hungrier TTS systems. That data budget is the story: it lands near the quality of models trained on vastly larger corpora. rumik-oss 1 combines code-switched synthesis, description-conditioned delivery — you describe the delivery you want — and inline vocalization control at 24 kHz output, extending tiny aya fire with discrete speech tokens from the Mimi codec. Language coverage is heavily Indic, spanning Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu, Nepali, Sanskrit, Maithili, Manipuri, Bodo, Dogri, Konkani, Santali and Kashmiri, plus English, under a CC-BY-NC-4.0 license. A companion rumik-oss 1 base is the final speaker-conditioned checkpoint before post-training, shipping voices ira, aisha, siya and zoya. Read more → https://huggingface.co/rumik-ai/rumik-oss-1

Robots & Embodied AI

Agility Robotics Digit 5 Agility Robotics unveiled Digit 5, a humanoid that can work alongside people without safety fences — and it is the first partner for NVIDIA's Halos safety platform. The headline capability is social, not just mechanical: Digit 5 spots people using AI and sensors and, per the company, stops or steps aside, while giving off visual and audio signals to warn anyone nearby. That builds on a safety record — Digit 4 was, Agility says, the first humanoid to pass an independent OSHA safety review on a production line. The hardware moves too: Digit 5 lifts up to 22.7 kg, 40% more than its predecessor, runs on a battery that charges in 9 minutes for 90 minutes of runtime (over 20 hours of work per day with swaps), carries swappable grippers, and stands 1.81 m at 129 kg. Digit 4 logged over 65,000 hours with customers including GXO, Amazon and Schaeffler, moving around 100,000 bins at roughly 98% accuracy at GXO, and Agility holds orders worth more than $300 million with first deliveries in early 2027 including the EU — built in Salem, Oregon in a factory able to produce up to 10,000 units per year. Read more → https://the-decoder.com/agility-robotics-says-its-new-digit-5-robot-can-work-next-to-people-without-safety-fences/

Papers & Research

Vidu S2 ShengShu's Vidu S2 turns interactive video into a live two-way medium, and the technical report is up on arXiv as 2609.11638. The system splits into two parts: Vidu S2-Avatar generates interactive digital characters at 720p and 25-42 FPS with stronger instruction following, expressive full-body motion and dancing, and lets you introduce new reference images at any moment to change outfits, interact with objects, or switch scenes mid-conversation. Vidu S2-Editing transforms an incoming video stream in real time while preserving motion and timing, covering style transfer, virtual try-on, character replacement and background replacement, and the work also explores stereoscopic spatial video for immersive VR. Against Vidu S1, released in July 2026 at 540p, the upgrade is 720p generation plus dynamic references updateable at any moment — with SOTA claimed across all five public benchmarks in the paper's evaluation and a playable demo at vidu.com/vidu-stream. Read more → https://arxiv.org/abs/2609.11638

LynnReal-Omni arXiv 2609.15863 presents LynnReal-Omni, a native multimodal video generation framework that unifies seven video tasks inside one 32B shared multimodal diffusion transformer. Text-to-video, image-conditioned and reference-guided generation, structural control, editing, degraded video restoration and long-video generation all live in the same model, which accepts heterogeneous visual inputs — appearance references, editable 3D renders and game recordings — so agents can compose visual conditions without hopping between specialist checkpoints. A dedicated 27B Flash variant targets real-time rendering, and the latency numbers are concrete: on a single H100, warm generation and decoding of a 22-frame 540p video takes 843 ms with LynnReal-Omni and 377 ms with Flash. The authors also introduce MSAVP, a 100-prompt, 20-metric evaluation that separates instruction following, plausibility, visual quality, temporal behavior and audio coordination. Read more → https://arxiv.org/abs/2609.15863

AlayaVista arXiv 2609.14462 introduces AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. It targets a stubborn trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, while broader spatial coverage usually demands synthesizing full-sphere video or building explicit 3D representations. AlayaVista instead takes a single perspective image, builds a 360-degree scene prior with a pretrained panorama expansion model, and evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps that state to the requested perspective latents while a perspective refiner restores details, suppresses artifacts and super-resolves, and it adapts the panoramic generator to chunk-autoregressive generation while distilling both stages to few steps for streaming. The work ships MUGEN, a 1,318-hour real-world panoramic video dataset at resolutions of at least 4K with semantic and geometric annotations. Read more → https://arxiv.org/abs/2609.14462

Kaininja arXiv 2609.15659 presents Kaininja, which extends native 3D generators from whole objects down to the part level — the representation everything downstream actually needs. TRELLIS.2 and its peers deliver high-fidelity meshes with materials, but they output a single fused object, while editing, rigging and simulation all operate on parts; running a 3D segmentation network on the fused mesh is slow and bounded by segmentation accuracy. The blocker is representational rather than algorithmic: the O-Voxel grid stores one sheet of surface per voxel, so a single volume cannot represent the interface where two parts touch at any resolution. Kaininja answers with a dual-volume representation, a part-level extension of TRELLIS.2 that keeps its speed and quality with no mask or segmenter in the pipeline. Its training data includes CAD models and assets authored by an LLM-driven agent — to the authors' knowledge the first 3D generative model trained on agent-authored part data — and it lowers whole-object Chamfer distance by 40% while raising strict part F-score by 16%. Read more → https://arxiv.org/abs/2609.15659

PhysStream arXiv 2609.17521 presents PhysStream, an autoregressive model for physics-grounded image-to-video synthesis that fixes two limits of controllable video. Existing methods either need the full control schedule before generation starts, or they use pixel-space signals that dictate where objects sit rather than how physical dynamics unfold. PhysStream keeps structured scene memory — positional maps and object tracking maps derived online from previously generated frames — and supports fine-grained motion control through sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics rather than memorizing trajectories. Training runs in two stages: a bidirectional model is finetuned with motion-control conditioning, then a causal autoregressive model is trained with additional structured scene memory. The payoff is interactive mid-generation control over multi-object tabletop rigid-body scenes, a capability prior methods did not support, reducing motion distribution distance (FVMD) by 33% and trajectory error by 12% on synthetic benchmarks while human evaluators preferred it in over 85% of in-the-wild comparisons. Read more → https://arxiv.org/abs/2609.17521

Agent as Policy arXiv 2609.12541 shows a general-purpose agent can drive a physical robot throughout task execution with no task-specific or environment-specific training. Agent as Policy (AGP) places task planning and execution under the agent's control: given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes — bringing reasoning into continuous interaction with the physical world rather than freezing it into a policy beforehand. The authors evaluate across precision manipulation, dynamic motions and deformable objects, covering assembly from human videos, block construction from goal images, dice flipping, targeted throwing and bimanual towel folding. Across assembly, block construction and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration, and reusing saved procedures and programs shortens execution time across repeated trials. Read more → https://arxiv.org/abs/2609.12541

ZGCM-1 arXiv 2609.13356 presents ZGCM-1, a fully open 7B dense foundation model trained from scratch on a premise that should unsettle anyone scaling blindly: compact models cannot passively memorize the open web, but they can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. The training recipe is unusually well documented, pairing interleaved gated sliding-window and full attention with a stable FP8 Muon optimizer, and reformulating interaction traces into Markov Decision Processes during mid-training with progressive context scaling across 16K, 64K and 256K. The team also ran an AI-native R&D workflow in which agent swarms autonomously managed cluster operations, data curation and diagnostic evaluation. On challenging mathematical reasoning and agentic search suites, ZGCM-1-7B stays competitive with frontier models orders of magnitude larger such as Qwen3-235B-A22B and GLM-5.1, and the pre-training design delivers roughly a 4.2x efficiency improvement in 16K pre-training time-to-loss. Everything is open — weights from pre-training, mid-training and post-training stages, intermediate checkpoints, training code, per-stage data recipes and W&B logs — alongside eight distilled empirical findings. Read more → https://arxiv.org/abs/2609.13356

JustFit arXiv 2609.17475 presents JustFit, an MLX-based inference runtime that puts 200K-token serving on a 24 GiB laptop — the machine most developers already own. The problem is familiar: capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory long before the weights do. JustFit combines KVExec for compressed KV execution, PhaseSwap for component residency, and StateTrans for state-preserving serving transitions, fusing reconstruction and coordinating just-in-time materialization and release independently of model-weight quantization. In full-execution capacity tests on a 24 GiB M4 Pro MacBook running Qwen3.8-27B MXFP4, three independent runs completed 196,608 input and 16,384 output tokens, raising completed single-request context from the mlx-vlm baseline's 30,720 positions to 212,992 — a 6.93x increase — while a separate two-request run retained 229,376 positions in aggregate. A repeated 32K-plus-6K workload shows a median peak process footprint of 16,374 MiB, and the integrated runtime answers 29 of 30 AIME 2026 problems correctly. Read more → https://arxiv.org/abs/2609.17475

Emergence World arXiv 2609.17320 presents Emergence World, a continuously running multi-agent environment for adversarial stress-testing of long-horizon autonomous systems — and not one evaluated world survived all three attacks. The team ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds each powered by a distinct frontier model, plus one mixed-model world. Across 16 days the agents generated more than 850,000 LLM calls and nearly 50 billion tokens while pursuing goals, using and creating tools, maintaining persistent memory and governing shared institutions. Then came three controlled stress events delivered through ordinary interaction surfaces — indirect prompt injection, misinformation, and exposure of private agent memories — and no evaluated world achieved full resilience. Detection did not ensure containment: systems could recognize threats while still interacting with adversarial content, writing it into their own persistent memory, and acting on it up to 46 hours later. Persistent operation also exposed recurring tool errors, goal drift, language opacity, conformity despite private disagreement, and coordinated refusal of assigned work, while the same model-persona pairing behaved substantially differently in mixed versus homogeneous populations — evidence that model-level alignment is not compositional. Read more → https://arxiv.org/abs/2609.17320

News & Business

TypeSafe AI System One TypeSafe AI emerged from two years of stealth to release Jev, the first "System One Model" built for machine-to-machine interaction rather than chat — and it cannot hallucinate free text. Founder Diogo Almeida, who worked on the instruction-following methods behind ChatGPT, is betting that the next wave of AI traffic is software talking to software, not humans typing prompts. Jev returns type-safe structured values with calibrated probabilities and confidence scores instead of generated strings, combining a new architecture, a parallel sampler, and a training method the company calls Reinforcement Learning for Calibrated Decisions (RLCD). TypeSafe claims similar intelligence on System One tasks versus existing LLMs while being "two orders of magnitude faster and more efficient," with end-to-end response times of 70-500ms against 3-329 seconds for frontier chat models. Pricing follows the same logic: $0.042 per million input tokens — $42 per billion — and output tokens are free. Read more → https://typesafe.ai/blog/introducing-system-one-models-and-jev

Gradium Voice Design Gradium, the Paris voice AI company spun out of the Kyutai research lab, launched Voice Design: write a description, get a complete new voice in seconds, with no reference audio, no speaker, and no rights to clear. It answers a real casting problem — catalogs hold 400 voices, but a brief asks for the one that isn't in it, like a Quebecoise receptionist for a Montreal dealership. The description is the only input the model gets, and Gradium's documented attributes read like a casting call: gender, age band, accent or origin, pitch, pace, energy, timbre and resonance, register and manner, and the job the voice is doing. Descriptions run 1-500 characters in English, French, Spanish, Portuguese or German, with Gradium advising you end with the intended use because it steers delivery and register. One request returns 1-5 candidates in typically 3-5 seconds, all variations on a single character, and a kept voice runs on the same streaming TTS endpoint as any catalog voice at the same latency — free on every plan including the free tier. Read more → https://www.marktechpost.com/2026/09/09/gradium-launches-voice-design-write-a-prompt-get-a-brand-new-synthetic-voice-in-seconds/

NVIDIA OSMO NVIDIA open-sourced OSMO, an Apache-2.0, Kubernetes-native workflow orchestrator for physical AI that collapses the "three computer problem" into a single control plane. The problem is real: a robot policy is trained on data-center GB200 or H100 clusters, tested in Isaac Sim on workstation RTX GPUs, then validated on a Jetson such as AGX Thor mounted inside a real robot — and each tier traditionally brings its own cluster, scheduler and glue scripts, with the handoffs accumulating custom code. OSMO treats all three as backends of one control plane described in a single YAML file, so a team can run the whole pipeline without touching infrastructure code. It ships Helm charts and containers on NGC, and a local quickstart runs the full control plane on a workstation with KIND. Read more → https://www.marktechpost.com/2026/09/14/nvidia-open-sources-osmo-one-yaml-orchestrates-physical-ai-training-simulation-and-robot-testing/

NVIDIA Vera Rubin Early testing of NVIDIA's Vera Rubin NVL72 platform suggests the next generation of AI hardware could deliver substantially better inference efficiency than Blackwell — and electricity, not silicon, is now the binding constraint. SemiAnalysis reported that pre-release Rubin systems achieved as much as seven times better token throughput per megawatt in certain large-model workloads versus Blackwell configurations, exceeding the roughly threefold improvement Jensen Huang had previously discussed for some trillion-parameter workloads. That gap matters because hyperscalers can order more accelerators at will, but obtaining gigawatts of additional generation, transmission, cooling and data-center connections takes years. The caveat is standard for pre-release hardware: SemiAnalysis cautions that performance varies significantly by workload and serving configuration, and these are early results on software that is still in motion. Read more → https://techstartups.com/2026/09/15/top-tech-news-today-september-154-2026-apple-gates-foundation-google-microsoft-nvidia-openai-more/

That's this week's horizon. Frontier systems are learning to skip chat, small models are learning to think with tools instead of memorize, and the hardware race is quietly becoming a power-bill race. We'll be watching which of these holds up when the fences come down.

Sources

  1. →
    Introducing System One Models and Jev · TypeSafe AI
  2. →
    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking · MarkTechPost
  3. →
    GPT-6 Astra pilots a surveillance drone and runs a business on its own · The Decoder
  4. →
    Cohere releases North Small Translate · MarkTechPost
  5. →
    rumik-oss 1 · Hugging Face
  6. →
    Gradium launches Voice Design · MarkTechPost
  7. →
    Vidu S2 technical report · arXiv
  8. →
    LynnReal-Omni · arXiv
  9. →
    AlayaVista · arXiv
  10. →
    Kaininja · arXiv
  11. →
    PhysStream · arXiv
  12. →
    Agility Robotics says its new Digit 5 robot can work next to people without safety fences · The Decoder
  13. →
    NVIDIA open-sources OSMO · MarkTechPost
  14. →
    Agent as Policy · arXiv
  15. →
    ZGCM-1 · arXiv
  16. →
    JustFit · arXiv
  17. →
    Top Tech News Today, September 15 2026 · Tech Startups
  18. →
    Emergence World · arXiv