Today skewed operator-heavy again, led by MiniMax H3’s 27.7x single-GB200 speedup, a sharp numerics warning from disaggregated inference, and a fresh wave of agent runtime and harness tooling.
NVIDIA SANA’s Sol Engine splits MiniMax H3 into a 4-step low-res draft plus 3-step LTX refinement pass, cutting single-GB200 latency from 414s to 14.93s and turning high-end video generation into a much more interactive workload.
The result is a concrete serving warning: 1,000 temp-0 requests produced 80 unique completions, with first divergence traced to token 103 because RMSNorm, matmul, and attention kernels were not batch-invariant.
Apodex 1.1 centers on async task decomposition, parallel agents, and continuous integration of findings across research, finance, and professional work, making the orchestration layer itself the shipped artifact.
A one-command converter from Claude Code sessions into Codex, Pi, OpenCode, and other harnesses points to a new portability layer for agent workflows instead of locking history into one runtime.
Google says its AI can forecast riverine floods up to 7 days ahead and urban flash floods up to 24 hours ahead, and the open-sourced hydrology framework turns that claim into something operators and researchers can inspect and build on.
Teneo now has setup guides for Claude Code, Codex, Manus, Cursor, Antigravity, VS Code, and Hermes, with plain-English queries for 50+ agents and per-call USDC settlement for paid queries via x402.
Farid Zakaria’s scheme stores ELF pieces in SQLite tables, sets the application ID to SELF, and can use binfmt_misc so Linux executes the database file directly.
The essay argues a malicious LLM could target the host machine running its weights by emitting token sequences that exploit the inference engine HTTP interface.
Commenters say Microsoft’s image tools add a unique identifier to each image, and one commenter argues that could reveal a meme creator’s identity through a subpoena.
Commenters say Agent Lightning v1.0.1 introduces an Agent Lightning Skill that uses a benchmark to improve prompts, tools, workflows, models, and reasoning settings.
AgentMercury builds persistent worlds with entities, services, tools, state, and executable cross-service invariants from high-level business scenarios.
The paper argues that tasks needing heterogeneous expertise, interdependent subtasks, parallel execution, independent verification, and persistent state exceed a single agent’s capacity.
SparsePR combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction to accelerate video transformers with training-free block-sparse attention.
Hydra-0 represents robot actions as pixel motion and reports 90.4% lower robot-motion error and 60.2% lower object-motion error than its action-conditioned baseline.
The paper defines infinite video editing, where a model generates the next segment of an ongoing stream instead of editing a fixed source clip frame by frame.
The study varies one factor at a time and finds on-policy distillation transfers reasoning behavior rather than answers, including from teachers that never solve the problem.
The talk describes a workspace that uses persistent instructions, read-only MCP data connections, and reusable analytical skills to produce Excel and PowerPoint deliverables.
The model ships in Hugging Face Transformers format and is slated for a hosted Qwen Cloud version with 1M context length by default and built-in tools.
The dataset release includes 1,440 de novo miniprotein binders against 16 targets, with binding calls, kinetics, raw sensorgrams, and report images from two vendors.
The benchmark compares memory systems on separate textual and coding tracks under a fixed evaluation contract that standardizes answer, eval, datasets, models, and configurations.
The dataset contains 57,937 quality-filtered traces from Qwen3.8-Max-Preview, GLM-5.2, and Kimi Code K3 across math, code, reasoning, tool-use, science, long-context, multilingual, and creative dialogue.
Apache Maka is a local-first agent workspace that runs tools under a sandbox boundary and records model messages and tool calls as recoverable execution facts.
Hermes Agent includes a built-in learning loop that turns experience into skills, searches past conversations, and can run from a $5 VPS to serverless infrastructure.
The repository collects official agent skills from teams including Anthropic, Google Labs, Vercel, Stripe, Cloudflare, Netlify, Trail of Bits, Sentry, Expo, Hugging Face, and Figma.
OpenHuman is a local-first AI system with persistent memory, orchestration, and research features, and the project says it was GitHub’s number-one trending repository for nine days after launch.