Today’s feed was dominated by agentic ML workflows turning into concrete products and benchmarks: NVIDIA pushed autoresearch from demo to playbook, Perplexity shipped wide research infrastructure, and model and safety claims kept colliding with real operator concerns.
NVIDIA moved beyond a demo and showed a full pattern: agents that set up environments, run RL loops, post-train models, and propose next experiments, with claimed jumps from 25% to 96.9% on star counting and to 93.35% on Cosmos 3 Nano in under a day.
Perplexity is exposing both the capability and the yardstick: Wide Research is now in the Agent API, and WANDR gives outsiders a benchmark for the deep-and-wide research behavior the product is optimized around.
LangChain’s tracing rollout and open coding harness point to the next layer of the agent stack: standardized session logs, reproducible coding runs, and shared telemetry across tools like Cursor, Copilot, Pi, and OpenCode.
New releases are selling context length, effort controls, and agent claims more than raw leaderboard wins, from GLM-5.2’s 1M-token window to Agents-A1’s 'trillion-parameter-level' pitch and Hy3’s product-tuned MoE update.
The gap between autonomous capability and safe deployment stayed visible: Cursor faced a long-lived disclosed 0day, Tailscale SSH surfaced a root-access bug, and new guardrails like destructive command hooks are arriving as practical countermeasures.
The post says the Technion paper found blindfolding people who provide imitation data produced substantially better test performance, including on shape insertion.
The post says distillation transferred traits such as Gemma 3 negative emotion into Qwen-base and Gemma 4 agentic misalignment into Nemotron Chat even after filtering prompts and rollouts that mentioned the trait.
The post proposes studying scalable oversight by training tiny models inside graphical abstractions of real-world problems and links code at AgoraForge.
The post argues that if models compute in superposition they need error correction that suppresses interference along specific feature directions rather than generic noise.
Simon Willison describes making a custom Codex Desktop pet by prompting GPT-5.6 Sol xhigh and using several rounds of gpt-image-2 to generate sprite assets.
The author says Juggler was built entirely solo in spare time, while commenters describe it as a GUI coding agent that keeps users involved instead of fully autonomous.
Commenters say the project explores JEPA-style world modeling for Mario, and one argues the example highlights long-horizon planning limits from chunking plans into intermediate goals.
Commenters describe a parallel search setup using 20 Codex accounts, large numbers of Lean 4 theorems, thousands of vCPUs, and proof embedding databases.
Commenters debate the framework, with one saying there are more than three loops and another arguing the tightest loop is inference, then tool use, then human oversight.
The paper proposes Direct-OPD, which runs RL on a smaller model and distills on-policy trajectories to a stronger model instead of distilling only the weak model's final policy.
The paper proposes PUST, which uses a lightweight proxy model for exploration so update signals can be generated asynchronously, reused, and transferred across models.
ABot-AgentOS adds a deliberative runtime layer for planning, memory, tool use, verification, and edge-cloud collaboration, and introduces EmbodiedWorldBench with 16 scenes and 200+ tasks.
ABot-N1 uses a slow-fast architecture that separates cognition from control to reduce coordinate drift and improve handling of long-tail semantics in visual language navigation.
The system curates 9.6K hours of egocentric pretraining data with 9x higher throughput than prior work and combines it with teleoperation, human-in-the-loop correction, and a world-model-enhanced VLA trainer.
Xiaomi-Robotics-U0 is a 38B-parameter multimodal autoregressive model that jointly trains text-to-image, image editing, embodied scene generation, and related embodied synthesis tasks.
The paper says modern LLM agents show myopic, polarized interaction patterns in multi-agent settings and proposes Multi-Agent Contextual Exploration (MACE) to probe peers and reduce regret.
LightMem-Ego is a streaming multimodal memory system that aligns egocentric visual and audio input on a shared timeline and organizes it into current, short-term, and long-term memory.
AdvancedMathBench introduces ProverBench, a proof-generation benchmark with 296 problems spanning undergraduate and doctoral qualifying levels instead of relying only on final-answer correctness.
The paper defines E-VQA, requiring models to output both an answer and spatiotemporal evidence including temporal segments and dense tracked object masklets, and introduces the ST-Evidence benchmark.
The guests describe Anthropic's developer platform as a three-layer stack of knowledge, execution, and coordination, with 'strategies' as meta-harnesses that assign different jobs to tokens.
This 0.9B audio model does single-pass transcription, diarization, timestamps, and acoustic event awareness for recordings up to 90 minutes across 50+ languages.
Audex-30B-A3B extends a 30B MoE text model with discrete audio tokens and an audio encoder to handle speech recognition, translation, TTS, audio generation, and speech-to-speech.
This dataset releases 4,665 Pi-compatible coding-agent trace sessions converted from 60 source sessions, with 3,799 tool actions and a median 2,365 characters of chain-of-thought.
EdgeBench evaluates agent learning over time on 134 real-world tasks, with 51 tasks and the full framework released publicly and 38,000+ hours of agent interaction analyzed.
This dataset supports Vera, a layered diffusion video editing framework that jointly generates an edit layer, alpha matte, and composite video to separate generated from preserved content.
Graphify maps code, docs, PDFs, images, and videos into a queryable knowledge graph, and its code parsing is fully local via tree-sitter AST without an LLM.
This repo packages 100+ open-source AI agents, agent skills, and RAG apps with end-to-end templates that work across Claude, Gemini, GPT, DeepSeek, Llama, and Qwen.
This repo offers small, composable agent skills intended for real engineering work and positioned as an alternative to more process-owning approaches like GSD, BMAD, and Spec-Kit.
Hallmark is a design skill for Claude Code, Cursor, and Codex that applies one of 20 themes, four verbs, 57 slop-test gates, and a pre-emit self-critique.
Vibe-Trading is a trading agent with API and MCP support, and the repo warns that a named X account, Virtuals project, and token contract are not official assets.
This proof-of-concept uses multiple agents for trading decisions and is being rebuilt into a persistent system with backtestable alpha models and optional live execution.
This Go project exposes OpenAI-style and Anthropic Messages-compatible APIs over pooled Grok Build, Grok Web, and Grok Console accounts with routing, quotas, and failover.
Go · ★ 5,892
Get this in your inbox
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.