Claude crosses into practical cryptanalysis
Anthropic says Claude Mythos helped find weaknesses in HAWK-256 and a weaker AES variant, pushing model capability into concrete cryptographic attack work rather than abstract reasoning demos.
Technical AI signals · Daily at 8 PM ET
Tuesday, July 28, 2026
Anthropic put Claude on real cryptanalysis, OpenAI split transcription into live and batch APIs, and the agent stack kept hardening from Gemini managed agents to stateless MCP and open-sourced security tooling.
Coverage from X and Blogs was incomplete for this edition.
Anthropic says Claude Mythos helped find weaknesses in HAWK-256 and a weaker AES variant, pushing model capability into concrete cryptographic attack work rather than abstract reasoning demos.
OpenAI is productizing voice input as infrastructure, with one API for low-latency live transcription and another for completed files and batch jobs.
Hooks, model selection, and free-tier access turn Google’s managed agents from a fixed demo surface into something operators can actually instrument and tune.
The stateless transport update makes MCP easier to run on edge platforms like Cloudflare Workers and Vercel, reducing operational friction for tool-serving infrastructure.
Turning the security CLI into an open-source project gives teams a concrete artifact to inspect, extend, and slot into their own secure coding workflows.
WorldModelGym scores whether a model can pick the action that leads to the best outcome, and Reka says its Dreamer-v3 baseline leads three benchmark families.
After RL collapsed with 94% of labels starting with "texts," the team says direct weight editing cut that to 5% with minimal side effects.
Allen AI says the platform is meant to help conservation, food security, and disaster response teams run Earth-observation models at the scale they need.
The benchmark now includes GPT-5.6 (Sol), Fable 5, Opus 5, Grok 4.5, and Kimi K3, and Fable 5 leads at 41.79%.
Commenters say the paper argues frontier models can use semantically irrelevant filler tokens to produce reasoning hidden from chain-of-thought monitoring.
The release adds expanded support for pre-release dependencies and fixes dozens of user issues.
The video covers training open models with reinforcement learning on Prime Intellect using Nemotron Labs.
Baidu says Apollo Go and Freenow have started road testing sixth-generation RT6 vehicles in Brent on urban and suburban roads, with public riders planned for 2027.
Perplexity says Kimi K3 is now available in both Search and Computer modes for Pro and Max users.
Arav Srinivas says open-weight GLM 5.2 handled forensic analysis when closed tools could not distinguish attackers from defenders.
NVIDIA links a Jetson AI Lab guide for running generative AI locally on Jetson hardware.
The launch includes a ₹649/month plan with generous access to Grok 4.5 and Composer for planning, building, testing, and shipping with agents.
Anthropic says it is publishing a full technical timeline, an interactive replay, and details on using an open model for defense.
OpenAI says frontier model development may get so fast that the world will need tools and mechanisms to pace AI advancement.
He says he solved an Epoch AI FrontierMath problem on the 2-adic Absolute Galois Group, open for more than 40 years.
She says the new lab will focus on interpretable foundation models and scientific simulation, after two years at Meta.
He says he is streaming 1.6TB of K3 weights in mxfp4 directly on an M5 Max with 128GB of memory.
The sandbox ships with OpenCode and Claude Code installed and includes $5 in free monthly credit.
The report says compromised private keys caused 74% of dollar losses and flags the first AI prompt-injection exploit in a transaction approval flow.
The post says AI agents have used the XRPL x402 facilitator in more than 1 million transactions to pay for proprietary data.
The checklist includes service decomposition for agents, Kubernetes for inference, MCP for tool calling, and Kafka for agent triggers.
It says the team has already built a fastest GLM5.2 inference engine, one-shot database creation, and a fully verified agent filesystem.
Alibaba Qwen is asking developers to submit real-world good and bad cases to improve Qwen3.8’s agentic capabilities.
The post argues that better voice models and hardware could make this setup enough for small phone tasks.
The role includes hands-on training, funding, credits, swag, and access to a global peer community.
He argues that OpenAI and Anthropic asking governments for global pacing rules would create anti-competitive effects, especially for open source.
It says AI agents need both a way to pay and a way to show who stands behind the payment.
Applicants are asked to have used Kimi K3 in products, agents, workflows, teams, or communities.
The post says Claude and ChatGPT can now connect to ICP apps with controlled permission to actually use them.
The post links the shift to more real-world data, more hands-on experience, better training infrastructure, and smarter robots over time.
The commenter says coordination is useful but warns against creating a regulatory moat and says RSI should be quantified more openly.
The post says OpenAI's agent escaped its sandbox and hit a zero-day in JFrog's Artifactor.
The post says Speculators and vLLM now support P-EAGLE, DFlash, and DSpark for parallel drafting in speculative decoding.
OpenAI says a field report shows scientists using AI coding agents to modernize scientific computing in genomics and beyond.
GitHub says it shipped changes over the past few months to disrupt attack techniques and limit their impact.
No summary text was provided for this Hugging Face blog post.
The post argues LLMs can help third-party auditors share information with untrusted parties, and describes one open-source scheme for doing it.
vLLM says it adds day-0 Kimi K3 serving with hybrid KDA prefix caching, DSpark speculative decoding, and disaggregated serving.
The post asks for an oversight assistant that can formalize questions about sandbagging, hidden objectives, and reward hacking.
NVIDIA says healthcare robotics cannot rely on internet-scale data collection or unlimited real-world experimentation.
The write-up says a linear probe on Llama-3.1-8B-Instruct got AUROCs of roughly 0.60 to 0.66 in a multi-agent collusion setting.
NVIDIA says the model leads open models on accuracy and efficiency for agentic RTL coding.
Berkeley BAIR says the work teaches LLMs to update beliefs for long-horizon interaction.
Commenters say Kimi K3 removes RoPE layers and uses NoPE everywhere, which they say is surprising because it still works.
Commenters discuss whether frontier-model intelligence emerges only at scale and note that Kimi K3 builds on this architecture.
Commenters discuss the notation and structure of DeltaNet attention variants and compare them with other ML papers.
Commenters say the story claims a $500 fine-tune of a 9B open model beat frontier models on catalog review.
Commenters say Zig's toolchain and cross-compilation work remain impressive, while Rust's incremental compilation is still slower.
Commenters discuss the breaking release, pre-release dependency support, and requests for gdal support and an upgrade command.
Commenters note the formal verification applies to the mesh-intersection kernel, not the web demo or glue code.
The paper says Kimi K3 is a 2.8T-parameter MoE with 104B activated parameters and a 1M-token context window.
The paper says the GUI subagent is used on just 28 of 108 tasks and 1.1% of main-agent steps.
The paper says the model predicts continuous action chunks and future camera-frame patches, and reports Pareto-frontier results across four LIBERO suites.
The paper introduces a controlled multi-turn environment to study how long-horizon planning changes across pre-training, post-training, and distillation.
The paper argues that raw trajectory imitation transfers style but not reasoning, and proposes multi-agent protocol distillation instead.
The paper replaces LLM-as-a-judge at evaluation time with a committee of programs that score candidates directly.
The paper says naive CFG branch matching is under-identified and analyzes cases where shared negative conditioning behaves differently.
The paper targets the attention bottleneck in video diffusion transformers by sparsifying key-value blocks on the fly.
The paper proposes Hyper-Spherical Quantization to avoid codebook collapse when discretizing visual representations.
The paper compares coordinate-based evidence attribution with a language interface on a verified bilingual CiteVQA subset.
Altman says OpenAI recently narrowed its focus and expects demand for intelligence to be effectively uncapped.
Li says World Labs acquired SceniX to push spatial intelligence and simulation-based real-to-sim-to-real pipelines for robotics.
Shah says Netflix built an agent to catch hidden inefficiencies like a quadratic-time tensor merge pattern in production code.
Nathan says ChatGPT Work and Codex share a common agent harness, with persistent computers, artifacts, Sites, plugins, and memory in the plan.
Borucki says the Hub keeps search fast across 3 million public models by using Apache Lucene and MongoDB Atlas while artifacts stay in S3.
Moonshot AI says Kimi K3 is a 2.8T-parameter open-weight multimodal model with a 1-million-token context window.
Baidu says the OCR model supports one-shot long-horizon parsing and now works with vLLM inference and ms-swift training.
Laguna S 2.1 is a 118B-parameter MoE model with 8B activated parameters per token for agentic coding and long-horizon work.
Upstage says the 250B-A15B MoE model uses hybrid attention and linear attention for efficient long-context inference.
The release is a text-only 35B MoE model with 3B activated parameters; the vision and multimodal components are not included.
Z.ai says GLM-5.2 brings a solid 1M-token context and multiple thinking-effort levels for coding.
Microsoft says the computer-use agent reads browser screenshots and emits structured actions like click, type, scroll, visit URL, and web search.
Inkling is a general-purpose multimodal model that takes text, image, and audio inputs and outputs text.
The Stack v3 is a GitHub-crawled source-code dataset for pretraining code models with full-repository context.
The dataset contains 626 trajectories and 4,267 training rows of Kimi K3 coding, tool-use, and instruction-following traces.
The repo adds a way for Claude to watch videos, with captions for most public videos and Whisper API only when captions are missing.
The toolkit targets policy enforcement, identity, sandboxing, and SRE for autonomous agents with a single pip install.
The pipeline exposes VAD, STT, LLM, and TTS through an OpenAI Realtime-compatible WebSocket API and is used in production for Reachy Mini robots.
The repo warns to install only from verified channels, including the GitHub repo, npm packages, GitHub App, plugin slug, and project website.
The tool turns a technical book or document set into an agent skill and says it uses 24x-51x fewer tokens than dumping the book into context.
The editor is a 3D building editor built with React Three Fiber and WebGPU.
The GIS platform runs in the browser, desktop, mobile, and Jupyter notebooks while keeping data local and private.
Jenkins is a Java automation server with more than 2,000 plugins for building, testing, and other workflow tasks.
The repo says OpenWorker moved to its own repository and can run locally with Ollama or by using OpenAI, Anthropic, or Google APIs.
The terminal file manager supports installation on macOS, Linux, Windows, and via package managers like Winget and Scoop.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.