Today was an operator-heavy day led by a new measurement of what coding agents actually read, another step-change in local MoE deployment, Blackwell kernel speedups, and AMD speculative decoding guidance in vLLM.
The clearest new agent artifact is empirical: 94K events across 557 coding sessions show instruction files and working notes drive 60.5% of reads, giving teams a concrete target for where to invest agent scaffolding and repo hygiene.
CAKE turned a kernel paper into deployable infra, upstreaming paired forward and backward kernels for SM100a and SM103a with up to 3.05x geomean speedups across 17 shapes on B200, B300, and GB300.
This is practical serving guidance rather than a benchmark screenshot: vLLM documents draft-and-verify, MTP, EAGLE-3, DFlash, DSpark, tuning knobs, and measured results for AMD deployments.
Local inference keeps widening fast: one project claims 39 tok/s for Qwen3.6 35B on an 8GB RTX 4060 laptop and 22-25 tok/s for DeepSeek-V4-Flash 284B on one 5090, while KTransformers keeps day-zero heterogeneous CPU-GPU support for new open releases.
The result is a useful warning for applied multi-agent systems: agreeable bargaining agents leak budget and concede too easily, while SocialRL lifts a 4B model to 0.627 average across six negotiation and scheduling games.
The Dreamer-4 world model was trained on 240 ABC-130k episodes across 48 tasks, and doubling the tokenizer from 96 to 192 latents improved reconstruction from 22.90 to 28.31 dB on one RTX 4090.
Open Design runs locally and can create web, desktop, and mobile prototypes, generate landing pages and dashboards, and export to HTML, PDF, PPTX, and MP4.
Commenters report the models found unpatched vulnerabilities and rooted the tablet, and one commenter says the Chinese models did so while the American ones fell back to safeguards.
FlowEvo compiles successful workflows into callable skills, stores them in a persistent bank, and tracks each skill's downstream utility at inference time.
WithEveryone generates group images with up to ten reference identities using addressed tokens, a structured identity-layout plan, and Layout-Grounded ID Loss.
TinyCast uses a zero-parameter spectral detector and has 146,505 parameters, placing it below 1.4M parameters among zero-shot entries declaring no test-data leakage.
NARU contains 1,481 questions grounded in 155 videos totaling 146.8 hours and evaluates narrative evolution and cultural understanding in Japanese long-form video.
The video says Z.ai released GLM-5.2, a 753-billion-parameter open-weights model, and claims the model is beating top Western models on major leaderboards.
The model card says Qwen3.8-27B is compatible with Transformers, vLLM, SGLang, and TokenSpeed, and that a hosted Qwen Cloud version with 1M context and built-in tools is coming soon.
The dataset release includes 1,440 de novo miniprotein binders against 16 targets designed by two Claude models and characterized by Adaptyv Bio and Twist Bioscience.
The platform is a unified evaluation system for long-term memory systems and memory-enabled agents, with the first verified leaderboard expected on August 12, 2026.
Ultra-FineWeb-L1 is an English Common Crawl corpus built with text extraction, language filtering, heuristic filtering, deduplication, and data cleaning.
The collection includes official skills from teams such as Anthropic, Google Labs, Vercel, Stripe, Cloudflare, Netlify, Trail of Bits, Sentry, Expo, Hugging Face, and Figma.
The tool turns a book, folder, or other source set into a skill and says that approach uses 24x-51x fewer tokens than dumping the book into context for one question.
The project warns to install ECC only from verified channels and says the install adds skills, agents, commands, and plugin-managed hooks to Claude Code.
OpenHuman is an early beta local-first system that claims better memory, orchestration, and tooling, and says it became GitHub's number one trending repository for nine days in a row.