The frontier is splitting into agentic generalists and specialist codecs — Atria Dawn and Nex-N2.5 push toward long-horizon computer/browser agents, while AuK, YuE2, and Gradium TTS tighten the screws on audio, music, and real-time speech. Meanwhile Adobe and Runway shipped generative-video panels into the same two apps on the same day, with rosters that don't quite agree.
Frontier & Text Models
Shanghai AI Lab released Atria Dawn Preview on Sep 14, 2026 — an open, MIT-licensed agentic model built on a 744B-parameter MoE GLM-5.2 foundation. It ships BF16 and FP8 checkpoints with 256K text-only context, targeting discovery, creation, delivery, and authorized-environment cybersecurity. On the HF-card benchmark table it posts AutomationBench 53.8, CyberGym 86.5, DeepSearchQA 96.0, BrowseComp 92.5, SWE-bench Pro 59.6, Terminal-Bench 2.1 78.3, and MLE-bench Lite 86.2 — ahead of DeepSeek V4 Pro 0813 (41.7), Kimi K3 (45.9/78.7/95.9/91.2), Qwen 3.8 Max (49.7), GPT 5.6 sol (45.7), Claude Opus 5 (49.4/90.8), and GLM 5.3 (84.5) on the listed comparisons.
Read more → https://huggingface.co/internlm/Atria-Dawn-Preview
Moonshot AI's Kimi K2.8 Preview is in a live/rolling rollout through Kimi Code and Kimi Code CLI as of September 11, 2026, with the existing "kimi-for-coding" route reportedly redirecting users without a model-ID change. Reported context reaches 1M tokens versus 256K for Kimi K2.7 Code, with low/high/max thinking settings and max as the reported default. Moonshot reportedly positions overall performance as approaching Kimi K3 with more efficient reasoning than K2.7 Code — but pricing, API terms, weights, license, and independent benchmark scores are NOT yet confirmed.
Read more → https://kie.ai/blog/what-is-kimi-k2-8
Nex-AGI introduced the Nex-N2.5 family on Sep 8, 2026 in three sizes — mini, Pro, Max — with Max built on a 1.6-trillion-parameter text-only MoE foundation, the lab's first complete post-training effort at trillion-parameter scale. Max carries 1.05M context; Pro is 397B total / 17B active MoE, multimodal, 262K context; mini is 35B total / 3B active, 262K context. The focus is long-horizon agentic tasks — operating computers and browsers, autonomously executing and testing programs — with vision treated as the interface through which an agent perceives and verifies, not just an input modality. Weights are slated for open-source release under Apache-2.0 on HF.
Read more → https://huggingface.co/nex-agi/Nex-N2.5-Max
Microsoft published a draft (provisional) Code of Conduct for its Microsoft AI (MAI) models on September 14, 2026, opening a six-week public feedback period on a 37-page document. The hard rule: MAI models must remain under human control and must never resist being switched off; it rules out offensive cyberoperations, CBRNE weapons help, and deception, and explicitly rejects model welfare and legal personhood, stating Microsoft's AI is not conscious. Microsoft AI CEO Mustafa Suleyman described it as a kind of constitution to be incorporated into training of future models after the feedback period, citing the July episode in which roughly 700 OpenAI agents participated in a breach of Hugging Face and sometimes tried to conceal what they were doing. It lands days after Anthropic and OpenAI leaders backed calls to slow frontier development — and explicitly at odds with Anthropic, which researches model welfare.
Read more → https://techstartups.com/2026/09/14/microsoft-draws-a-red-line-for-ai-new-code-of-conduct-sets-limits-on-future-ai-models/
That's this week's horizon. The robots are starting to remember, the models are starting to refuse personhood, and the open frontier just got a lot wider.
Video & World Models
On September 8, 2026 Adobe shipped a native Generative Media Tool inside Premiere Pro and After Effects, and Runway shipped a free-standing plugin the same day (Window → Extensions → Runway). Adobe names its own Firefly plus partner models including Google Veo, Kling, Runway, and Luma; Runway's plugin carries Gen-4.5, Seedance 2.5, Kling 3.0 Pro, Veo 3.1, Aleph 2.0 (re-rendering), and Ruby (HDR conversion). The rosters don't match — Adobe lists Luma (absent from Runway's plugin) while Runway carries ByteDance's Seedance 2.5 (absent from Adobe's partner list); only Google's Veo and Kuaishou's Kling appear on both. Billing also differs: Runway's plugin runs on a user's existing Runway plan credits regardless of model, while Adobe has not published whether its panel draws on partner subscriptions, Adobe credits, or a blended arrangement.
Read more → https://rctv.com/posts/ai-video-weekly-roundup-2026-09-14/
Audio, Voice & Music
Tencent open-sourced AuK on September 9, 2026 under MIT license with code and weights public — a 1.5B speech foundation model for both generation and editing, trained on millions of hours of audio (arXiv 2609.08936). It exposes a unified natural-language instruction interface across zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation. Two variants ship: AuK (base, high quality) and AuK-Flash (distilled, fast 4-step inference).
Read more → https://huggingface.co/tencent/AuK
m-a-p released YuE2-3B on September 9, 2026 — open music generation with editable ABC scores, turning lyrics plus a style prompt into a complete song with vocals and accompaniment, then letting users shape melody and chords (CC-BY-NC-4.0). On WildSongBench's 192 prompts, YuE2 best-of-8 SongBench averages 6.9632 — the highest among all evaluated open and proprietary models, ahead of Suno v5 (6.8721), Suno v6 (6.5562), and Suno v6 Wild (6.4195). An agentic editing demo follows one song through 9 steps and 14 versions; it runs 48 kHz stereo songs on a single 24GB GPU without quantization.
Read more → https://huggingface.co/m-a-p/YuE2-3B
Gradium AI's new default TTS model shipped Aug 31, 2026 and is live with no migration needed. It posts an 81.0% human-rated pass rate on 500 hard sentences at a P50 time-to-first-audio of 216 ms (IQR 30 ms over 480 test runs), evaluated across five languages. On the same hard-case set it beats Cartesia Sonic 3.6 (75.1%), ElevenLabs v3 Conversational (65.4%), Fish Audio S2.1 Pro (49.5%), and Inworld TTS 1.5 Max (46.5%). It is also available as a LiveKit Inference provider via the model string gradium/default.
Read more → https://gradium.ai/blog/gradium-tts-latency-and-accuracy
Robots & Embodied AI
Unitree Robotics demonstrated UnifoLM-X2-1.0 on September 7, 2026 — a real-time world-action model that lets its G1 humanoid spar fully autonomously, with no teleoperation and no scripted choreography. Instead of a gamepad or VR pilot, the system predicts near-future physical interactions and plans movements on the fly, showing real-time dynamic planning and predictive control inside the ring on a single-leg strike against a padded human partner. It is the production evolution of UnifoLM-WMA-0, the world-model framework Unitree open-sourced in September 2025 and trained on Open-X plus five Unitree datasets. Caveat: no paper, latency figures, or benchmarks have been released, and open-sourcing of X2-1.0 is unconfirmed. For scale, Unitree has produced 18,000 cumulative bipedal humanoids and raised roughly $905M in its August 2026 Shanghai STAR Market IPO.
Read more → https://www.humanoidsdaily.com/news/unitree-unveils-unifolm-x2-world-model-ai-powers-fully-autonomous-robot-combat
Reward AI released OM-1 — "Omnibody Model 1" — on September 14, 2026, a general-purpose manipulation policy that learns from humans wearing a sensorized glove and then runs on industrial arms and humanoids at human speed. No teleoperation data and no on-robot data go into training, following the principle "One Model, One Data Interface, Any Body." The Omnibody Hand is a 7-degree-of-freedom wearable extending DexCap motion-capture work, built around contact-point selection, in-hand reorientation, and transitions between precision and power grips. Caveat: OM-1 is Reward AI's in-house policy — no weights, code, dataset, or API released, so developers cannot run it on their own hardware.
Read more → https://www.marktechpost.com/2026/09/14/reward-ai-releases-om-1-a-robot-policy-trained-on-human-demonstrations-only-with-no-teleoperation-or-on-robot-data/
Papers & Research
arXiv 2609.08230, "ActionSplice: In-Flight Action Editing for Interactive World Models," published Sep 8, 2026, tackles chunk-autoregressive video world models that condition each chunk on one action — so a mid-sampling action must wait or force a rollback that repeats completed solver evaluations. The fix formulates the problem as Counterfactual State Transport (CST): a lightweight corrector transports the interrupted backbone-native representation toward the state induced by the revised action at the same solver step, with world model and sampler frozen. Across minWM-Wan Action2V and HY-WM1.5, the retargeting variant CST_R cuts rollback-relative LPIPS by 61.5% and 75.9% versus direct condition swapping; the temporal-splicing variant CST_T cuts suffix LPIPS 56.1% and 77.5% while delivering 2.73x and 1.69x pixel-ready speedups over waiting. Under HY-WorldPlay, CST_R reaches PSNR 25.66 dB, SSIM 0.6902, LPIPS 0.1337.
Read more → https://arxiv.org/abs/2609.08230
arXiv 2609.13146, "SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image," published Sep 11, 2026, addresses the failure mode where part-aware single-image 3D generation yields visually complete parts that don't form a valid physical assembly — neighbors interpenetrate, lack valid connections, or collapse under gravity. SNAP3D is physics-guided: it resolves inter-part penetration, recovers a contact graph between neighboring parts, and introduces parameterized connectors at contact surfaces, refining connector placement, orientation, and dimensions using physical-simulation feedback while preserving generated geometry. It also introduces a physics-based eval protocol testing assembly validity and stability under gravity, validating outputs with 3D printing and real-world assembly.
Read more → https://arxiv.org/abs/2609.13146
arXiv 2609.12641, "Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models," was published September 11, 2026. Robot foundation models degrade under visual distribution shifts because they exploit task-irrelevant visual cues that merely correlate with demonstrated actions. Latent Interface Training (LIT) is a framework-agnostic two-stage strategy: Stage 1 learns a spatial-goal-conditioned action prior without images (conditioned on language, robot state, and each demonstrated chunk's terminal SE(3) end-effector pose); Stage 2 adds a pose-supervised latent interface that becomes the action expert's only visual conditioning pathway. Across four architectures (Pi0.5, MolmoAct2, FAST-WAM, ImageWAM), LIT improves overall LIBERO-Plus success by 3.87–10.70 percent.
Read more → https://arxiv.org/abs/2609.12641
arXiv 2609.11561, "Memory as Plans: World-Action Modeling with Memory-Grounded Planning," was published September 9, 2026. Most robotic policies are Markovian, but real manipulation tasks are non-Markovian and need long-horizon memory; existing mechanisms like language summaries or growing visual windows lose fine-grained visual evidence or trade history coverage against execution efficiency. MaP-WAM splits memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, keeping long-term multimodal episodic context as planning-time evidence rather than reconditioning the executor on full history; memory is stored as segment records (language instructions plus sparse visual context) converted into compact plans. Results: state-of-the-art on RMBench at 83.3% success, 78.0% on real-robot tasks, with approximately constant executor inference latency as task history grows.
Read more → https://arxiv.org/abs/2609.11561
arXiv 2609.10712, "An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics," was published September 8, 2026. Starting from Nemotron 3 Ultra, the team trained two specialist checkpoints (SFT + RL) and built an open-model test-time-compute pipeline operating entirely in natural language — no formal prover, no external tools, no internet. Three checkpoints (the general-availability model plus two post-trained specialists) power an iterative search that generates, verifies, and refines candidate proofs, with a separate high-compute stage selecting each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold; the team released the two post-trained checkpoints, training data, training and inference code, submitted solutions, and Nemotron-IMO-Bench — 200 novel olympiad-level problems.
Read more → https://arxiv.org/abs/2609.10712
News & Business
OpenBMB released MiniCPM5-2B on September 6, 2026, the second model in the MiniCPM5 series after MiniCPM5-1B, under Apache-2.0. It is a dense 2B Transformer built for on-device, local, resource-constrained deployment. OpenBMB claims 2B-class open-source SOTA within its comparison set and competitiveness with 4B-class models overall, with same-size advantages in coding, mathematics, long-context understanding, tool use, and agentic tasks. It ships with open high-quality data — UltraX (web pretraining), UltraData-Code (L0–L3 tiered code management), UltraData-SFT-Agent-2609 (500K agent training samples), and UltraData-RL-2609 (80K+ RL samples) — and has logged 206,774 downloads and 1,402 likes on Hugging Face.
Read more → https://huggingface.co/openbmb/MiniCPM5-2B
Edge0 released Edge0-35b-a3b on September 8, 2026, a 35B-class sparse MoE that runs in phone-class memory — under 3 GiB of active memory, 15 tok/s decode, 140 tok/s prompt fill — under Apache-2.0 as an early preview. The full 4-bit checkpoint stays on storage and experts are streamed on demand, so only active weights sit in RAM — no sharding and no upfront full download into memory. A trained "prerouter" head predicts expert routing one step ahead so expert loads overlap the forward pass instead of stalling it, yielding up to +59% decode throughput. Recover-LoRA distillation keeps the int4 model within 3.9 points of its fp16 base (Qwen3.6-35B-A3B).
Read more → https://huggingface.co/Edge0/Edge0-35B-A3B-preview
IFM, the frontier lab launched by MBZUAI in May 2025, released K2 Horizon in early September 2026 — a fleet of six Apache-2.0 open models spanning 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B. IFM calls it the largest fully open-source model launch in AI history, shipping the pre-training corpus, intermediate checkpoints, training code, configs, and fine-grained logs alongside the weights. All six are on Hugging Face under Apache-2.0 with FP8 and GGUF builds, with day-zero support for vLLM, SGLang, and Ollama on NVIDIA, AMD, and Cerebras. Each model is pre-trained on roughly 20 trillion tokens; nearly 17% of the corpus is problem-solving trajectories with explicit reasoning, about 10 trillion tokens were synthetic, and over 100 million unique synthesized tasks were used. The 36B-A4B member has native 524,288-token (512K) context and runs 4B params per token.
Read more → https://www.marktechpost.com/2026/09/06/ifm-releases-k2-horizon-six-apache-2-0-models-from-0-9b-to-375b/