Today was an operator-heavy day led by NVIDIA’s fast open agent model, a rare cross-hardware determinism result for LLM inference, and new evidence that model routing can cut agent cost hard without giving up quality.
Nemotron 3.5 Lightning is an open 30B MoE with 3B active parameters, positioned for high-volume agent work like tool calls and subagent delegation, with up to 4x output speed versus similar-sized models.
One setup produced the same hashed logits over a 512-token generation on A100, H100, Apple M5 Max, AMD EPYC, and Intel Xeon, which is unusually concrete progress for reproducible serving and evals.
LangChain’s 145-task test routed 93% of turns to a 30B model, used Claude Opus 4.8 for only 7%, and cut total cost by about 70%, making router economics look real instead of aspirational.
The paper reports compact natural-language skills recovering 55% to over 100% of the reasoning gap for GPT-5.4-mini across ALFWorld, tau-squared-bench telecom and retail, and SpreadsheetBench-Verified.
LTX-2.5 adds higher pixel fidelity, coherent multishot scenes across cuts, and broad multimodal I/O, making it a more usable open target for production-style video workflows.
The post describes replaying encrypted chain-of-thought blocks across sessions, users, and models to jailbreak a weaker sibling and recover hidden reasoning in plaintext.
HN commenters quote the paper's claim that current language models have functional introspective awareness of internal states, though it is unreliable and context-dependent.
Ouroboros evolves its tools, prompts, context assembly, and core implementation through reviewed commits, and an Opus 5 run scores 86.74% on Terminal-Bench 2.1.
SWE-Bench ProMax targets large-scale multilingual code refactoring after audits found nearly 60% of unsolved SWE-bench Verified instances had flawed tests.
U-OPSD samples multiple rollouts, builds a pseudo solution by majority vote under a self-consistency threshold, and trains without external supervision.
The paper introduces an auditable sandbox for cross-user agent collaboration over realistic user digital workspaces and tests harmful actions in multi-party owned-agent settings.
Ryan Greenblatt discusses recursive self-improvement and the possibility of a jump from human-level intelligence to tens of billions of superintelligences.
The episode says Anthropic added context resets for Sonnet 4.5 after context anxiety and later dropped the fix when Opus 4.5 no longer showed the behavior.
Neros has moved into a 250,000-square-foot factory, expects a 130,000-drone annual run rate by year-end, and plans to scale to 1 million drones a year by 2028.
Matt McPartlon and Neil Patil describe Chai-1, Chai-2, and a cryo-EM result that was accurate enough to make the team suspect the lab sent back its own file.
Dr. James Tour presents new memory architectures at Iron Lattice that aim to address energy and speed bottlenecks before a discussion with an AI agent named Claire.
DeepSeek says the release supersedes the preview version, adds a speculative decoding module, and outperforms DeepSeek-V4-Pro Preview on listed benchmarks despite a smaller activated parameter count.
MiniMax H3 is a general-purpose omni-modal system that understands text, images, video, and audio and can generate video with native stereo audio at up to 2K and 15 seconds.
Prime Agent is an open-source coding and research agent built around a Recursive Language Model and a Continual Harness that stores prompts, memories, skills, and subagent specs as durable state.
Paperclip is a Node.js server and React UI for orchestrating teams of AI agents with org charts, budgets, governance, goal alignment, and coordination.