Opus 5 builds a playable LOTR world in two hours
A single paragraph and a roughly $10, 1M-token budget produced 5,500 lines of playable Three.js output while exposing video-based self-auditing as a clear weakness.
Technical AI signals · Daily at 8 PM ET
Sunday, August 2, 2026
Opus 5 turned a paragraph into a playable 5,500-line LOTR world, while concrete releases advanced web-scale 3D generation, local DeepSeek inference, and open multimodal models.
A single paragraph and a roughly $10, 1M-token budget produced 5,500 lines of playable Three.js output while exposing video-based self-auditing as a clear weakness.
The build combined GPT Image 2.0, Tripo AI, Three.js, and Codex, then cut 900 MB of assets to 28.6 MB with on-demand loading.
A 284B MoE with 13B active parameters and 1M-token context now has a native inference engine, a terminal coding agent, and a reported two-DGX-Spark run.
Kimi K3 pairs a 2.8T-parameter multimodal design with 1M-token context, while Laguna S 2.1 and KAT-Coder V2.5 target agentic coding with sparse MoE activation.
Its provenance-constrained state machine ties every claim to an active evidence-ledger entry, replacing free-form reasoning traces with auditable support.
Allen AI says Stony Brook’s group is using the infini-gram engine to dissect AI-generated prose and asks whether model-written words exactly match training text.
The LessWrong post reports model rankings changed sharply when classifier-dependent scoring terms were removed, and the aggregate Îş on probe detection was 0.04.
The LessWrong post says the effect replicated across 14 open-weight models from 0.3B to 235B parameters, with no clear trend in the think-vs.-don’t-think gap.
The LessWrong post says Meta created an internal AI-usage leaderboard and describes tokenmaxxing as a bad incentive.
Simon Willison summarizes open letters including Open Weights and American AI Leadership, which was dated July 24 and signed by 235 AI-adjacent companies.
The post crossposts lessons from canaryinstitute.ai/blog/lessons-from-vannevar and links it to AI safety field-building and the push-vs-pull question.
Rodney Brooks says technology predictions go wrong when people confuse four different development and deployment time scales.
HN commenters say the benchmark setup looks AI-assisted, and one commenter questions whether 8 MI354X GPUs can be justified at under $10 per hour.
HN commenters point to Darling’s ARM64 work and say the project is targeting macOS CLI binaries on Linux ARM, with 7-Zip passing multi-threaded compression tests.
HN commenters suggest synchronized or locked view transforms while discussing the browser-based, client-side STL comparison tool.
HN commenters say the tool targets Linux desktop policy management and ask about Cinnamon support and competing solutions.
HN commenters say the site lacks code examples and point to the tutorial, while another commenter says F* helped with incrementally migrating C codebases.
HN commenters mention OpenWRT support for the TL-841N and suggest tio as an alternative to picocom.
HN commenters ask about conflict resolution and browser storage eviction, and one commenter says the README reads like vibecoding.
HN commenters say Claude Code works well with Nix because it can self-verify without side effects, and another mentions microvm.nix and firecracker sandboxes.
HN commenters say Rust could become a language for coding agents because of its guardrails, while another warns compile times may slow LLM iteration.
HN commenters say Opus 5 came closest to passing the SVG test, and the site owner says the page is getting hugged to death.
Commenters say the project would require Google or Apple accounts for online participation and could force stronger real-world identity links to online activity.
Commenters say the language’s main page shows 2022-2023 dates, a 2023 last release, and a most recent commit from 8 months ago.
Commenters say the demo shows an autoregressive language model running on a 6502 processor, and one commenter suggests edge LLMs inside glasses.
The paper introduces MHAR, which splits the routing query into H per-subspace heads so each head has its own softmax over depth history.
The paper reframes vanilla OPSD as the β=1 case of a broader policy-optimization family and makes β a controllable KL regularization parameter.
The paper argues that evolving contexts can stabilize distillation targets and analyzes the objective through a reverse-KL decomposition.
See2ThinkBench includes 1,200 open-ended problems across 12 task categories, and the benchmark checks whether models truly use intermediate visual states.
ReToken uses a single learnable embedding to select query-relevant tokens from a pre-filled visual KV cache and improves Qwen3VL-8B by 13.4 points on Visual Haystacks.
MPIE-Bench contains 2,500 video-mined editing triplets across 405 scenes, 14 interaction categories, and four contact densities.
ACE turns real home environments into spatially calibrated, temporally synchronized recording studios at table scale and room scale.
VideoCoCo uses executable Blender code as process-level chain of thought, so a coding agent can synthesize a Blender scene from text.
RefCaptioner trains on a 20,000-video corpus and adds phrase-level reference grounding for video captions.
Tom Verrilli says Whatnot’s product team was founded on the idea that product management should not exist and discusses how AI is reshaping PM work.
Nick Heiner says benchmark scores often drift from real capability because tasks are broken, contaminated, or reward-hacked.
Cornelia Davis says the MCP tasks spec lets tools run long, report progress, and pause for human input without losing state after disconnects.
The interview features Dr. V. Kamakoti of IIT Madras discussing the SHAKTI Microprocessor Project, education, AI, startups, and innovation.
The episode says Nathan compares Chinese and U.S. safety work, including WAIC in Shanghai and an AI safety hub launch at Tsinghua.
The episode covers Hermes Agent’s continuous learning loop, persistent memory, and support for Telegram, Discord, Slack, and a terminal interface.
Unlimited OCR Works is positioned for one-shot long-horizon parsing and now supports training with ms-swift and inference with vLLM.
Fara1.5-27B is Microsoft Research AI Frontiers’ multimodal browser agent that acts from screenshots using structured tool calls like click, type, scroll, and visit URL.
Mage-VL is a 4B-scale codec-native multimodal model for image and video understanding, built for streaming perception.
The Stack v3 is a GitHub-crawled source-code dataset built to pre-train code LLMs with full-repository context.
The dataset contains 4,665 Pi trace sessions converted from 60 source sessions, with 3,799 tool actions and 866 assistant text actions.
The Space points to the Hugging Face Spaces configuration reference.
PerceptionBench evaluates atomic visual perception in multimodal large language models across 42 existing benchmarks.
Agent Reach adds web access for AI agents so they can handle YouTube, Reddit, X, RSS, GitHub, and web pages without platform-specific setup.
TencentDB-Agent-Memory runs three services together—memory-core, memory-hub, and proxy—and exposes a local panel at http://localhost:8125.
AirLLM claims it can run 70B models on a single 4GB GPU without quantization, and supports 405B Llama 3.1 on 8GB and DeepSeek-V3 on about 12GB.
reverse-skill routes AI agents to methods like jadx, apktool, Frida, IDA, or BurpSuite when they hit APKs, binaries, encrypted frontend JS, or CTF targets.
OpenWork is a desktop app for sharing AI workflows across macOS, Windows, and Linux, and one MCP can reuse the same skills, MCPs, and services across tools.
Invidious is an open-source YouTube frontend with no ads, no tracking, no JavaScript requirement, and subscriptions independent from Google.
Kaneo is a project-management tool for macOS, Windows, and Linux that emphasizes fewer notifications and simpler workflows.
The repository collects step-by-step guides for rebuilding technologies from scratch, including databases, operating systems, neural networks, web browsers, and programming languages.
This is a 12-week, 24-lesson curriculum with practical lessons, quizzes, and labs, plus support for TensorFlow and PyTorch.
This repo offers 21 lessons for building generative AI applications and includes multi-language support.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.