Kimi turned its K3 launch into a full open stack drop, NVIDIA and Microsoft both pushed benchmarked agent harness stories, and Washington may be moving frontier-model access into the prerelease lane.
Coverage update: Some Blogs sources were unavailable for this edition.
Moonshot didn’t just post a model card: it says K3 brings 2.8T MoE scale, native vision, 1M context, open weights, and a supporting release of MoE comms, attention kernels, agent runtime, and day-0 vLLM serving.
NVIDIA’s NOOA pitch and GitHub’s harness essay make the same operational point from different ends: model performance now depends heavily on the runtime that manages context, tools, and workflow execution.
Microsoft says MAI-Cyber-1-Flash plus MDASH hits 96% on CyberGym at half the cost, while operators are already reframing the market in price-performance terms with claims around Grok 4.5 versus Sol, Opus 5, and Kimi K3.
NVIDIA paired alliance politics with task-specific results, saying it is contributing open agent research to OSAA while Nemotron 3 Ultra reached a 97.1% pass rate on agentic RTL coding.
A reported plan would give federal agencies up to 30 days of access to frontier models before release, pushing prerelease testing closer to a formal national-security review process.
MSR’s AI Foundations for Power Grids workshop at NeurIPS 2026 is accepting submissions until August 29 on benchmarks, model training, reliability, foundation models, and deployment.
Goodfire says spiky-sounding and round-sounding words land on opposite ends of a specific activation direction in Llama and Gemma, independent of meaning.
Together AI says Sol leads pass@1 on 904 DeepSWE rollouts, while Kimi K3 wins pass@4 at 2.8x the solves per dollar and routing between them reaches about 85.6%.
Simon Willison summarizes an updated guide that shifts from chat use to agentic systems and says Gemini still lacks an established entry in the Codex/ChatGPT Work/Cowork category.
This paper studies the scaling behavior of training a transformer-based vision-language model from scratch on multimodal inputs under a fixed compute budget.
The paper proposes SPA, a two-stage method that fits a spectrum prior from training data and uses it at inference time to correct intermediate predictions.
Epsilon says its Bengaluru team processes over 400 billion consumer actions daily and that generative AI cut planning-to-activation cycles from 12 weeks to 2.
Moonshot says Kimi K2.7 Code is a coding-focused MoE model with 1T total parameters, 32B activated parameters, and about 30% less thinking-token use than Kimi K2.6.
Microsoft Research says Fara1.5-27B is a multimodal computer-use agent for browsers that acts from screenshots with tool calls like click, type, scroll, visit URL, and web search.
This repo gives Claude the ability to watch video by using yt-dlp and ffmpeg, with captions for most public videos and Whisper only when no captions exist.
MediaCrawler is a multi-platform scraping tool that says it can extract public data from sites like Xiaohongshu, Douyin, Bilibili, Weibo, Tieba, and Zhihu.