The frontier turned two dials at once today: down on price and sideways into the interface. Anthropic shipped its cheapest small model yet and cut Sonnet cache reads in half, while OpenAI pushed GPT-6 into an interactive UI and dumped 372 machine-authored math results onto GitHub. Below that line the robotics money kept moving — Mecka raised $60 million, Nvidia circled another $1 billion for Figure, and Boston Dynamics handed the wheel to Amazon's former Alexa chief. Even the research papers pointed the same way: the smartest way to fake intelligence is getting cheap, and the hardest thing to fake is still physics. Here is the edition.
Frontier & Text Models

Anthropic's Claude Haiku 5.5 is its cheapest, fastest and most capable small model yet — about 75% cheaper to run than Haiku 4.5. It is built for high-volume, cost-sensitive work like summaries, classification, database queries and compactions, and pairs as a subagent under Opus 5.5 and Sonnet 5.5 on coding runs. Anthropic reports 72.4% on OSWorld 2.1 (offline subset), 45.9% on Humanity's Last Exam without tools and 57.4% with tools, and 39.2% on Terminal-Bench 4.0 — and it is the first Haiku-class model with an adjustable effort setting. Alongside it, Anthropic halved Claude Sonnet 5.5's cache-read price (about 20% cheaper on most agentic work) and added a monthly API credit for Max and Team subscribers.
Read more → https://www.anthropic.com/claude-haiku-5-5

OpenAI's GPT-6 rollout reshapes ChatGPT itself: answers arrive as interactive interfaces instead of plain text. The new "Intelligent UI" adapts the format to the question — comparisons render side by side, explanations become charts, and users can have small tools built inside the chat, such as a savings calculator or a game. GPT-6 can also respond while it is still thinking, which OpenAI says cuts wait times by 44% and, in internal tests, beat GPT-5.6 on difficult web searches. Paying users get GPT-6 Sol; free and Go tiers get GPT-6 Luna one day after the paid global rollout began on October 7.

OpenAI published 372 mathematical results produced by an internal frontier model, and put them on GitHub rather than in journals. Each result is meant to solve an open problem or make substantial progress toward one, including improvements to major computer algorithms and advances tied to the Riemann hypothesis. Nearly every result came from a single prompt to a single agent, at roughly three hours of ChatGPT Pro compute on average — a sharp contrast to the company's earlier Navier-Stokes proof, which needed a swarm of 10,000 agents. Many proofs ship with formalizations in Lean, the machine-checkable proof language, which OpenAI frames as a way to ease the peer-review bottleneck.

OpenAI launched a Decisions API that turns messy evaluations into a yes, a no, or a pick-one. The endpoint evaluates text, images or both and runs about ten times faster than the Responses API, in public beta with general availability coming. So far it supports only gpt-6-luna, priced at $0.10 per million input tokens with output tokens free, and it returns one of three shapes: yes/no probabilities, a pick from predefined categories, or a scale-based rating — aimed at uses like damage detection, inquiry routing and document classification, with zero-data retention and HIPAA-compliant use in the US and Europe. OpenAI also cut its paid API tiers from five to three, with monthly usage limits of $500, $5,000 and $200,000.
Robots & Embodied AI

Mecka raised a $60 million Series B led by Sequoia to scale the human-motion data behind robot training. Nvidia, Microsoft's M12, Qualcomm Ventures and Samsung joined as new investors in the October 7 round, which funds expanded data collection, research and commercial robot deployment. Mecka builds capture hardware and processes recordings into structured motion and 3D data, and says its enterprise offering also covers robot integration, on-site collection, model training and ongoing operation. The company says it passed a $100 million revenue run rate in June and projects $300 million by year-end — company-reported figures, not audited results.
Read more → https://www.humanoidsdaily.com/news/mecka-60-million-series-b-robot-training-data

Nvidia has discussed investing another $1 billion in Figure AI, The Information reports, as the humanoid developer raises at roughly a $38 billion pre-money valuation. Figure is pursuing large-scale compute and training-data projects, including a multibillion-dollar Nscale compute agreement and its Index effort to collect human video for robot training. Nvidia is already an investor: in September 2025 Figure announced more than $1 billion in committed Series C capital with Nvidia among the backers, at a $39 billion post-money valuation. The report does not establish that a new deal has been agreed.
Read more → https://www.humanoidsdaily.com/news/nvidia-figure-billion-investment-talks-report

Boston Dynamics named Rohit Prasad, formerly Amazon's AI chief, as CEO effective October 7. Prasad spent 12 years at Amazon as senior vice president and head scientist for Alexa and Artificial General Intelligence, helping build Alexa and later leading the Nova foundation-model family; before that he spent nearly 14 years at Raytheon BBN Technologies. The hire is framed around bringing advanced AI into robots for industrial and commercial settings, spanning Spot, Stretch and the Atlas humanoid. He succeeds Robert Playter, a 32-year company veteran who retired earlier this year, and is expected to join the board subject to approval.
Read more → https://www.humanoidsdaily.com/news/boston-dynamics-rohit-prasad-ceo

Deep Robotics tested its IP66-rated DR02 humanoid on traffic duty in the rain in Hangzhou's West Lake district. In an October 6 video, the robot handles navigation, obstacle avoidance, traffic-violation detection and pedestrian interaction — asking a rider to move behind the stop line, announcing a green light, reminding riders to wear helmets and answering questions about local food, with on-screen labels showing pedestrian and vehicle avoidance. Deep Robotics presents the work as testing for future support of traffic officers, not routine deployment. The full-body IP66 rating is the point: the demonstration links the robot's claimed dust and water protection to a real street-level task.
Read more → https://www.humanoidsdaily.com/news/deep-robotics-dr02-rain-traffic-test-ip66
Papers & Research
![]()
A new benchmark tests whether video world models actually obey physics — and the best of eight still falls short. "World Models' Last Exam in Physics" is a measurement-based benchmark of 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism and surface tension, each paired with an initial image, a generation prompt and predefined physical criteria. Because it measures observable physical relationships directly rather than judging against reference videos, it does not require a human or model to grade the output. Across eight video generation models and 1,280 videos, the researchers found persistent physical inconsistencies and large variation between tasks.
Read more → https://arxiv.org/abs/2610.08791
![]()
DepthWorld gives robot world models the 3D geometry that RGB-only video rollouts lack. Current video-based world models can look correct frame by frame yet fail to compose into a consistent 3D world, which breaks their use for policy evaluation, improvement and planning. The authors built a calibration pipeline that fuses learned stereo depth with a joint factor graph to recover a robot's shared kinematic parameters and per-scene extrinsics — applied to the DROID dataset it yields DROID-3D, with dense metric depth and recalibrated multi-view extrinsics hitting under 0.7-pixel reprojection error on 90% of episodes. DepthWorld then trains a Stable Video Diffusion-based world model that jointly preserves geometry and appearance.
Read more → https://arxiv.org/abs/2610.08780
![]()
nanoMuse is an open-source answer to Meta's Muse: a personal agent that lives on the devices you already own. The report defines the personal agent through five questions and three horizons, then reconstructs how Meta's Muse works from Meta's public record and a copy of its production prompt, marking each statement by source. Its open counterpart, released under GPL-3.0, puts one agent on every device a person owns — with hands on the phone screen and the computer — sharing one conversation over a relay anyone can run. Every action passes through a "Sentinel," memory is stored as files the person can read, and the underlying model is the user's choice; size and cost are given as estimates.
Read more → https://arxiv.org/abs/2610.08699
![]()
MIMESIS trains a user simulator that behaves less like a helpful assistant and more like a real person. Interactive agent training usually leans on off-the-shelf assistant LLMs, whose helpfulness makes them overly cooperative, explicit and behaviorally homogeneous compared with real users. MIMESIS is purpose-built on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns, and its 9B model reaches a SOUL-Index of 65.7 — above the strongest frontier model. Against Claude Opus 5, the strongest baseline on RealUserSim and SimulatorArena, it improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points; the team then freezes the simulator and trains an agent against it.
Read more → https://arxiv.org/abs/2610.09484
![]()
A new paper shows why "keep the better score" quietly degrades self-improving LLM loops. When a model rewrites its own instructions and keeps whatever scores higher on a small evaluation set, it is doing selection under measurement noise, and reused eval sets turn that noise into the winner's curse. In runs where Qwen models rewrite their own instructions — every candidate also scored on 600 held-out items — most proposals after the first were harmful, and the final selection-set score exceeded held-out accuracy by 13 to 20 points with 16 selection items, shrinking to 1 to 5 points with 256. The tested acceptance rules did not beat greedy acceptance over whole runs, a caution for anyone scaling self-improvement on a thin benchmark.
Read more → https://arxiv.org/abs/2610.09239
![]()
Sensor-Language-Action models use language as the bridge between raw sensor data and action. Existing sensor models largely stop at perception — recognizing states or predicting outcomes — leaving action modeled separately in narrow, task-specific label spaces. The SLA framework connects multimodal sensor observations, natural language and actions in one model, letting heterogeneous actions be represented, predicted and explained while staying grounded in sensor evidence. To back it, the authors assembled a large-scale benchmark spanning more than 116,000 individuals, 79 sensor modalities and 60 action groups, with a captioning pipeline that aligns user context, sensor dynamics and action evidence, then built OpenSLA for hierarchical action prediction, state understanding and action explanation.
Read more → https://arxiv.org/abs/2610.08244
News & Business
ElevenLabs says India is now its second-largest market, with 100 million conversations run through its voice agents in a year. The company says more than 250 Indian businesses and enterprises now use its technology and more than 25,000 Indian developers and businesses hit its API monthly — about 2.5 times the figure a year earlier. Over 70% of those conversations were in Hindi, Kannada, Tamil and Telugu; the Eleven v4 model supports 14 Indian languages, with a goal of all 22. It signed an MoU with the Karnataka Innovation and Technology Society, while CARS24 runs more than 400 voice agents that have automated three million minutes of calls.

Google launched Playground, a browser-based tool that turns text prompts into playable games. The service, built with Unity, lets US adults create and tweak game rules, physics, characters and environments through back-and-forth prompting, then test and share the results — running on Gemini, Nano Banana and Lyria. Finished games can be shared by link or published to a public gallery, with some genres supporting leaderboards and multiplayer; the tool is free, though Google One subscribers get higher weekly limits. Google and Unity also announced Unity Spark, a professional-grade tool with access to the Unity Asset Store, in closed beta later in 2026.

Liquid AI released two open-weight multimodal "decision models" that return typed answers in one forward pass — with zero output tokens. d1-3B reads text and images; the smaller d1-omni-600M reads text with an image, or text with audio. Neither writes prose: given a state and a set of named questions, each returns a calibrated probability for every allowed answer, which makes them far cheaper to run than a generative model whose output must be parsed. Both checkpoints are on Hugging Face, load through Transformers and have day-one llama.cpp support, targeting NVIDIA DGX servers, RTX workstations and Jetson edge boards, under the LFM Open License v1.0 with free commercial use below $10 million in annual revenue.

Zuckerberg and Chan's Biohub is coordinating a $1.8 billion push to build AI models that predict cell behavior. The nonprofit had already pledged $500 million in April for its five-year "Virtual Biology Initiative"; now Meta, Google DeepMind and Isomorphic Labs are contributing a combined $300 million, and the US Department of Energy is investing over $500 million in lab measurements and compute over five years. The NIH is coordinating datasets built with more than $500 million in prior federal funding, which Biohub will standardize for AI training. Commercial funders get one year of exclusive access to data they paid for before it goes public; a first dataset is expected in about a year.

SpaceX is seeking roughly $40 billion to buy Nvidia AI chips, turning the AI buildout into industrial-scale debt finance. Bloomberg and the Financial Times report the package would include about $10 billion in bank loans and roughly $30 billion in investment-grade debt, with Apollo Global Management expected to lead and distribute it to investors and Pimco reported to have held early talks. The transaction is expected to close in 2027. The deeper story is structural: AI developers increasingly need financing for hardware whose upfront cost runs into tens of billions before the compute earns a dollar, pulling banks, private-credit firms and bond investors into an industry once funded mainly by tech balance sheets and venture capital.
That's this week's horizon. Cheaper small models, interfaces that think while they answer, and robots financed like infrastructure — the frontier's real story is no longer a single benchmark, but how fast it gets cheap and how well it stays honest about physics. Until tomorrow.