Today was an operator-heavy day led by NVIDIA’s concrete skill-evaluation results, Ornith-1.5’s benchmark-first open-model push, and a fresh wave of harness-native agent infrastructure for cost, safety, and training.
NVIDIA’s clearest artifact today is a verified-skills benchmark stack: 300+ skills tested on real tasks, with reported gains of 41 points in correctness, 39 in effectiveness, and 35 in efficiency, plus a concrete method for measuring how much context a skill actually saves.
Ornith-1.5 is notable because it ships as 9B dense, 35B MoE, and 397B MoE while claiming 86.1 on Terminal-Bench 2.1 and 86 on SWE-Bench Verified, making it one of the day’s most concrete open entries in the coding-agent model race.
Vral’s Router is a practical operator move: pick the model per request, keep output quality flat, and cut spend by about 40%, which is exactly the control plane teams want as model portfolios get wider.
The harness layer kept shipping today: Apache Maka reached incubating status with a burst of open-source activity, while OneCLI surfaced as an OSS sandboxed team harness, reinforcing that deployable agent infrastructure is hardening below the model layer.
LEGO-RL, Agent Lightning v1.0, and Agentic ESOpt all push the same concrete direction: train long-horizon coding agents through the harness they actually run in, instead of rewriting the agent loop for research-only RL setups.
NVIDIA FLARE is being used to build federated workflows for vision-language models such as visual question answering, captioning, and image-text reasoning.
The transcript argues that coding agents can make lines of code a productivity indicator because software engineers once faced a hard upper limit on output.
Commenters say the paper shows model reasoning traces can reflect implicit biases and still produce answers that do not match the written chain of thought.
FreeToken treats a personal machine as an elastic inference platform and co-designs model loading, expert residency, CPU-GPU execution, state reuse, and memory management.
The study evaluates 26 performance metrics across dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and context mechanisms.
The paper studies when skills help by testing representation, outcome annotation, retrieval difficulty, and cross-framework robustness across harnesses and LLMs.
Jared Joselowitz says Dora has made about 200,000 real clinical calls across 20 UK hospitals and is contracted to reach a million patients within two years.
The panel covers AI projects in academic libraries, including workforce development, collections and infrastructure, AI literacy, ethics, and campus partnerships.
The episode describes a Pocket Analyst Tool pattern that uses Model Context Protocol gateways and deterministic Python sandboxes to compress exploratory analysis into sub-minute runs.
The dataset includes 1,440 de novo miniprotein binders against 16 targets designed by two Claude models and measured by two contract research organizations.
The leaderboard evaluates long-term memory systems under one contract with fixed Answer, Eval, datasets, models, configurations, result review, and publication.
This Hugging Face repository provides FP8-quantized weights with block size 128 and says the artifacts are compatible with Transformers, vLLM, and SGLang.
OpenViking stores memories, resources, and skills as a virtual filesystem under the viking:// protocol and loads content in three tiers: L0, L1, and L2.
Munder Difflin wraps tools such as Claude Code, OpenAI Codex, Grok, Qwen, and GitHub Copilot CLI into a multi-agent harness that runs on a user's machine.
Superpowers is a coding-agent methodology built from composable skills and starter instructions for tools including Claude Code, Cursor, and Gemini CLI.
Shell · ★ 274,248
Get this in your inbox
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.