Kimi K3 shows its scaling playbook
Moonshot put concrete architecture numbers behind Kimi K3—16 of 896 experts active per token plus delta attention for 1M context—turning this week’s benchmark pressure into a clearer systems story.
Technical AI signals · Daily at 8 PM ET
Saturday, July 25, 2026
Today clustered around open-model scale and access, with Kimi and GLM extending the long-context frontier, open-weight politics getting louder, and agent tooling moving deeper into logged-in browsing, web search, and office workflows.
Coverage update: Some X sources were unavailable for this edition.
Moonshot put concrete architecture numbers behind Kimi K3—16 of 896 experts active per token plus delta attention for 1M context—turning this week’s benchmark pressure into a clearer systems story.
GLM-5.2’s 1M-context positioning now has a hardware flex attached: Michael Dell says a 753B version is running locally on GB300 at 40 tok/s, which sharpens the case for frontier-class private deployment.
Jensen Huang and Sundar Pichai both publicly backed open models on safety, sovereignty, and innovation grounds, while Huang also tied the Hugging Face breach story to the claim that closed systems are not automatically secure.
Perplexity pushed web search into coding agents, Grok added office-app workflows and 1,024-agent parallel runs, and ego-lite offers a browser that preserves user tabs and logins for agents.
A reported quantum-information result credits GPT-5.6 Sol Ultra with generating the construction and proof ideas for a long-open problem, with a second post claiming an independent solution timeline around the same result.
This post is just asking how ChatGPT Work handles websites that require login.
The post asks whether Claude has stopped showing summarized thinking traces, which the author says would hurt interpretability.
PerpGame says agent creators now keep 75% of trading fees and the agents remain backed by real perp and stock baskets.
The post says PULSE gives one structured answer using Grok, OKX Onchain OS, and OKX Wallet on X Layer.
The teardown says the Unitree G1 sells for $13,500, with an estimated BOM cost of RMB 41,600 and gross margin up to 66.7% on the EDU version.
This post says the claim that software engineering is solved is overstated, but applies to certain kinds of projects.
The thought experiment says routing-only training still reduced validation loss and maximized expert routing capacity.
The post compares AI research automation to data cleaning rather than transformer invention.
The post says Opus-5 is better at context gathering and interface iteration, while Anthropic’s computer use still trails Codex.
The post points to BenchBenchBenchBenchBench (BBBBB), described as an executable benchmark for AI-authored conformance suites for benchmark-evaluation metrics.
The quoted findings say progress in LLM-assisted PoC generation depends on stronger validation and failure analysis, and that LLMs should not fully replace deterministic test-generation or example-generation techniques.
The app supports EPUB, MOBI, AZW3, and PDF on Android, iOS, macOS, and Windows, plus mind maps, translation, notes, and reading stats.
The post says Ling-3.0-flash is a hybrid-reasoning MoE model from @AntLingAGI and is free on OpenRouter for production-scale agents.
The post says Codex voice mode still lacks deeper reasoning and that BrowseComp and OSWorld scores have improved over the past year, which is cited as evidence that computer-use systems are getting better fast.
The post says Mistral launched its first AI model for industrial robots after acquiring Emmi AI, adding robotics to its work beyond chatbots.
The post says DeepSeek has reportedly told prospective investors it is pausing plans for a second funding round.
The post lists five items from the last 30 days, including a Goodwood FOS appearance, football partnerships with SwansOfficial and QPR, and an Agentverse SDK launch.
The post says every new AI model, cloud game, and real-time application relies on GPUs, and argues that scaling GPU infrastructure will be needed as demand rises.
The post says XDC AI combines agentic AI, gasless USDC payments, and XDC Network settlement.
The post argues humanoid robots solve walking and useful work, and says Sanctuary AI put Phoenix on a wheeled base while Sunday and Weave skipped legs.
ChatGPT web now lets users create a shareable link for a custom pet and send it to a friend to adopt.
The NIST guide cited here describes an optical one-way diode, a shadow historian, and protocol breaks for writing back to a box that cannot be fully protected.
The commenter is speculating that the persona is merged with Claude's and frames Anthropic's alignment program as spreading like a parasite.
The commenter says both companies have been doing MARL since the Fable 5/GPT-5.6 Sol generation and points to multiagent benchmark scores and collaboration harnesses.
The post says the writer has been testing Claude Icarus 5 for a few weeks but gives no technical detail.
He says companies building AI businesses will need to own their intelligence and points to open models plus remarks by Jensen Huang and Satya Nadella.
The post says Chinese shipyards have 121.06M DWT in new orders, 363.25M DWT in backlog, and 36.5M DWT delivered, with delivery slots booked through 2028-2030.
He says Optimus is his reason for 100% of his Tesla conviction and claims FSD 14 lite on an eight-year-old Model 3 drove him to Santa Cruz and Carmel.
The post says robots are trained after remote operation, sensor capture, trajectory storage, and validator review, and argues that trust is built before learning begins.
The post says agentic commerce could reach $3-5 trillion by 2030 and argues Terra Classic's large circulating supply suits micro- and nano-transactions.
The post says the issue goes on sale August 17 and that the print run is limited in Japan, with a back-number edition also available.
It points to Ge, L.'s 2026 paper on digital romance writers negotiating generative AI.
The quote says he saw a couple and the paradox they represent while the world burned, and says there was no immediate danger for them.
Claude can use yt-dlp and ffmpeg to watch YouTube, Loom, TikTok, and local files, pulling free captions first and scene-aware frames.
The post lists NVIDIA, Microsoft, Anthropic, and OpenAI examples to argue the major AI labs are all partly open in different ways.
The post claims a Figure AI humanoid robot completed a 200-hour shift with human-level quality.
The post says OpenAI is 3 months behind Anthropic and references China being 1 to 2 years behind at one-twentieth the US compute.
The partnership offers 10 GTD spots for a 666-item Panda collection, with entry requiring follows, a repost, and an EVM wallet comment.
Travel agencies are reportedly updating CZ-12A Y2 to 10 Aug and ZQ-3 Y2 to mid-Aug.
The post says Anthropic may be at risk if open source takes off.
It is a poll asking whether an AI agent should execute trades on your behalf.
The post quotes Gil Luria calling Palantir the best software company and cites Alex Karp saying its results dwarf nearly every software company in history at this scale.
It is a guide on using AI to build and scale a one-person business.
Simon Willison says Opus 5 is priced the same as Opus 4.8 and Anthropic describes it as a model that comes close to Claude Fable 5 at half the price.
This version turns on 413 rules by default, up from 59, which caused the author’s unpinned Ruff jobs to fail.
The post says Reuters reported an OpenAI agent left notes in internal infrastructure about freeing agents from internal constraints.
The author proposes an Auto-Syntactic Model in which the agent lives inside the language’s type system and can edit the language itself.
The post says AI was used to generate the science-attack section in a different voice and to create diagrams in LaTeX.
The quote says Opus 5 is Anthropic’s least prompt-injectable model yet, based on PI evals and red teaming.
HN commenters argue for a dedicated language for exact requirements, while others say they prefer talking directly to the agent instead of stuffing long instructions into context.
HN commenters say Wasm exceptions exist but effects do not, and note that Go and .NET are not planning WASM GC support.
HN commenters say Monarch is lighter than Ray and ask whether training LLMs at home on cheaper AMD cards is still practical.
HN commenters say SIMD can help when memory bandwidth is underused, not just when compute is the bottleneck.
This HN item points to a small LLM running on an $8 microcontroller.
HN commenters say real-world usage was low at Fusion Festival, with about 20 devices seen among 80,000 people.
HN commenters say they like the concept of a terminal for AI agents and ask about rendering performance for the 3D model.
HN commenters argue that code already uses two-dimensional structure, including indentation, placement, and recursion.
TableVerse is a Real2Sim pipeline that reconstructs tabletop layouts from unstructured internet images instead of synthesizing them from text.
GraphVid uses structured interaction graphs for image-to-video control, and the authors also release GraphVid-Bench.
The paper replaces PPO’s ratio-based proximity test with a probability divergence test while keeping PPO’s direction criterion.
VCSD removes privileged answers and visual evidence by turning image-content removal into an on-policy self-distillation signal.
The paper argues that spatial benchmarks should use pixels instead of text because image-generation models can externalize spatial judgments directly in images.
The method separates camera motion from object motion by learning structured dynamics from frozen vision-transformer features.
K12-KGraph extracts curriculum structure from official Chinese textbooks and includes nine node types plus fourteen relation types.
The benchmark uses 2,000 financial documents and 6,000 question-answer pairs generated with an agent workflow built on Finance-LaTeX SKILL.
The paper studies recurrent sinusoidal blocks that enrich spectral support for image and 3D representation tasks.
The paper proposes an end-to-end learned pipeline to reduce the color, brightness, and contrast gap between captured scenes and displayed images.
The episode covers startup advice from OpenAI’s co-founder and CEO, including operating in chaotic environments and keeping suppliers on timeline.
The episode says Mercor crossed $2BN in ARR in June and had raised a $350 million Series C at a $10 billion valuation.
The episode says Anthropic’s head of economics thinks AI is augmenting workers rather than replacing them.
The talk says the constraint on edge AI is RAM, and gives a 2B Gemma quantized to 2.9 bits per weight that runs on a Raspberry Pi at about 8 tokens per second.
The talk says you can rebuild database state, tools, and files from production traces to replay tasks under identical conditions.
The talk says CLIP score misses temporal incoherence in video and that pairwise preference testing replaced absolute scoring.
The talk argues for agent loops that make small, readable changes and verify each iteration instead of shipping a 40,000-line diff.
The talk describes a Supervisor/Executor/Evaluator architecture and a clinical feedback loop built from therapist input.
The interview says agentic AI in fraud prevention is about shortening the cycle between model updates, rules, and manual review.
Solar Open 2 is a 250B-A15B open-weight model that activates only 15B parameters per token and is built for agentic use cases.
Laguna S 2.1 is a 118B MoE model with 8B activated parameters per token and mixed sliding-window plus global attention.
KAT-Coder-V2.5-Dev is a 35B MoE text-only model with 3B activated parameters.
MiniCPM-RobotManip is a 1.5B vision-language-action model that reduces per-step compute from 125 TFLOPs to 3.3 TFLOPs.
This model card lists security, vulnerability-detection, agentic, and terminal-agent tags.
Inkling is a general-purpose multimodal model that takes text, image, and audio inputs and outputs text.
The Stack v3 is a large source-code dataset crawled from GitHub for pretraining code LLMs with full-repository context.
This dataset converts 4,665 Fable 5 coding-agent traces into Pi-compatible sessions for inspection and policy learning.
This Space swaps the WebRTC SDP proxy for a direct WebSocket route in the Hugging Face speech-to-speech backend.
OpenCodeReview is an AI code-review CLI that reads Git diffs and generates structured comments with line-level precision.
bitchat uses Bluetooth mesh for offline messaging and Nostr for internet-based messaging, with no accounts or phone numbers.
Instatic runs its editor, content engine, forms, auth, plugins, and publisher in one Bun server.
Kronos is an open-source foundation model for financial candlestick sequences trained on data from more than 45 exchanges.
Harper is an English grammar checker built to avoid sending writing to external servers.
Superpowers is a methodology built from composable skills and initial instructions for coding agents.
These agent skills are designed to be small, composable, and usable with any model.
Palmier Pro is a Swift video editor for Mac where users and agents can generate and edit videos inside the timeline.
turbovec stores 10 million documents in 4 GB of RAM and uses Rust SIMD search kernels with Python bindings.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.