OpenAI pushes Codex into vuln repair
OpenAI is positioning Codex Security as a real code-security workflow, with GPT-5.6 Sol aimed at finding, validating, and fixing vulnerabilities instead of just suggesting patches.
Technical AI signals · Daily at 8 PM ET
Friday, July 17, 2026
Today was about operationalizing agents: OpenAI pushed Codex deeper into vulnerability repair, uv added OSV-backed malware blocking, GitHub opened up Copilot’s engine as an SDK, and NVIDIA expanded both its fine-tuning and embedding stack.
OpenAI is positioning Codex Security as a real code-security workflow, with GPT-5.6 Sol aimed at finding, validating, and fixing vulnerabilities instead of just suggesting patches.
`UV_MALWARE_CHECK=1` turns package install security into a defaultable control point by checking OSV before remote installs and blocking known malware.
GitHub is productizing the Copilot runtime across six languages, giving teams a standard engine for planning, tool use, and file edits inside their own apps and workflows.
NeMo AutoModel now reaches Diffusers training while Nemotron 3 Embed is being pushed as a leaderboard leader, extending NVIDIA’s pitch from infrastructure into adaptation and retrieval layers.
The conversation is shifting from abstract agent risk to system design: sandbox boundaries, permissioning, and classic security controls are becoming first-class parts of agent deployment.
The paper introduces error routing that trains strictly positive and negative neurons without backprop and reports image recognition, locomotion, and Craftax results.
The post reports Inkling early-late CKA around 0.8 versus around 0.5 elsewhere, and says Laguna XS 2.1 in bf16 and nvfp4 had nearly identical J-space after quantization.
Perplexity Agent API now supports custom skills.
The post says one engineer may be "10x" with Claude while the rest of the organization has not adopted it the same way.
The clip is presented as uncut, no-speedup execution of a single end-to-end policy assembling parts.
The repost describes a drone made to appear like a thin background blur by spinning the whole craft at high speed rather than hiding it with paint.
The workflow blocks motion in Unreal Engine with Cartwheel Swing, then re-renders with Seedance 2.0 so character performances can be edited and rerun.
The prompt asks Ultra mode to do a thorough review of all `unsafe`, UB, and related concerns in a crate.
The post recommends building mockups, schemas, data models, and proofs of concept first to avoid wasted token spend.
The post says builders converge on "superpods/supernodes" and cites ganging up 128 cabinets from a single 950 SuperPoD design.
The post updates a paper and says its TAC travel-booking benchmark now appears in the UK AI Security Institute's Inspect Evals with a leaderboard at compassionbench.com/tac.
GitHub frames a decision model for AI-era code changes around code getting cheaper to write while ownership cost does not.
The post proposes benchmarking some subjective conceptual tasks by asking models to predict a specified person's judgment instead of treating the task as having a single objective answer.
The post organizes alignment techniques across dimensions including internals versus outputs, SFT versus RL, and online training versus toy-domain training.
The experiment used LoRA fine-tuning on open-source models in the Tinker API with GPT-5.5, Opus 4.7 and 4.8, Gemini 3.5 Flash, and Kimi K2.6 as agents.
Simon Willison quotes Kimi K3 replying, "Is there something I can actually help you with today?" after refusing to leak its system prompt.
Commenters cite OpenRouter figures saying open models went from 40% to 63% share in four months and aggregate tokens rose from 888B to 4.19T.
Commenters argue AI-generated submissions and AI judging can break competition evaluation, including claims of prompt-injected winner declarations.
Commenters say the demo is an encrypted ResNet-20 on CIFAR that claims under 200ms latency and about 3x over recent state of the art, while another commenter reports a caveat after inspecting the JS and network logs.
Commenters point to SQLite `.expert` for index recommendations and one commenter cites a tool, `uvx s3-credentials`, for creating AWS backup credentials.
Commenters say the live site was quickly flooded with spam, and one links a related project called `honeyprompt` that uses LLMs to craft responses across multiple protocols.
A commenter asks whether the linked URL contains any content.
VideoChat3 is presented as a fully open video MLLM that aims to improve generalization across video types while releasing training code, strategy, and datasets for reproducibility.
SEED converts completed on-policy trajectories into hindsight skills and distills them back into the policy to fill the gap between sparse episode rewards and token-level learning.
SearchOS turns search progress into explicit shared state to reduce repetitive loops and wasted search budgets in single- and multi-agent web search.
LongStraw targets million-token RL post-training under a fixed GPU budget and says inference contexts are approaching 2M tokens while post-training often stays at 256K or below.
RoboTTT scales robot visuomotor context to 8K timesteps, described as three orders of magnitude beyond prior policies, without increasing inference latency.
BadWAM studies World-Action Drift Attacks, where small visual perturbations break alignment between a world-action model's predicted future and executed action.
GRASP trains agentic RAG systems with RL over semantic search, keyword search, and paragraph-reading actions so they can choose retrieval timing, method, and granularity.
The paper proposes Subspace-Aligned Rewiring, a post-hoc edit that keeps the spectral component of RL updates while removing orthogonal components to preserve reasoning gains and reduce interference.
The study frames on-policy distillation as an exploration catalyst rather than a capability ceiling increase, and identifies Student-Teacher Mismatch and another guidance-quality pathology as failure modes.
The paper argues interactive world models should be structured around a recurrent action-state-observation loop to preserve rules, long-horizon consequences, and real-time interaction.
This Latent.Space episode is about Kimi K3 and describes it as a 2.8T-A50B open model priced at Sonnet 5 levels.
CoreWeave SVP Corey Sanders discusses AI-native infrastructure, GPU optimization, training and inference workloads, and agentic development.
Hard Fork covers Apple's accusation that OpenAI sought hardware trade secrets, discusses OpenAI's Sol and Anthropic's Fable access change, and interviews Erik Brynjolfsson on AI and jobs.
Inkling is an open-weights multimodal model that takes text, image, and audio inputs and generates text for agentic, tool-use, coding, and RAG applications.
GLM-5.2 is presented as a flagship long-horizon model with a 1M-token context plus coding with multiple thinking effort levels.
Hy3 is a 295B-parameter MoE model with 21B active parameters and 3.8B MTP layer parameters, with Tencent saying it was post-trained using feedback from 50+ products.
This 0.9B model handles transcription, diarization, timestamps, and acoustic events across 50+ languages in a single pass for recordings up to 90 minutes.
OvisOCR2 is a 0.8B page-level document parser that outputs Markdown in reading order for text, formulas, tables, and visual regions, with a reported 96.58 OmniDocBench score.
Needle is a 26M-parameter pure-attention encoder-decoder distilled from Gemini 3.1, with claimed production speeds of 6000 tokens/sec prefill and 1200 decode.
Agents-A1 claims trillion-parameter-class performance from a 35B agent model, and the repository notes a new 4B release plus quantized variants.
UltraX publishes five English pretraining corpora of about 20B tokens each, using a refinement model to predict structured insert, delete, and modify edits that are then executed deterministically.
The dataset contains 4,665 Pi trace sessions converted from 60 source sessions, with 3,799 tool actions, 866 assistant text actions, and median reasoning length of 2,365 characters.
This Hugging Face Space points to the Spaces configuration reference.
PostHog pitches an open-source product stack with a self-driving mode that turns product signals like errors, rage clicks, and failed queries into reports and pull requests.
Open Interpreter says it reimplemented the provider-recommended Kimi Code harness in Rust and can switch harnesses with `/harness` for low-cost models.
The tool builds a Tree-sitter structural map, tracks changes incrementally, and feeds graph-aware MCP context so coding agents read only the changed parts they need.
The workshop materials include model sweeps over quality-per-dollar and quality-per-second, plus a 400-line prompt decomposed into skills, code execution, and callable agents.
Hallmark is a design skill for Claude Code, Cursor, and Codex with 20 themes, 4 verbs, 57 slop-test gates, and a pre-emit self-critique step.
turbovec says a 10 million document corpus drops from 31 GB as float32 to 4 GB, with online ingest and SIMD search that beats FAISS by 10–19% on ARM in cited configs.
The demo runs Bonsai 1-bit and ternary models locally across Metal, CUDA, Vulkan, ROCm, or CPU, and the new 27B line adds vision, OpenAI-style tool calls, MCP servers, and reasoning modes.
Recent releases add selective removal of one failed document from a knowledge base and multimodal image extraction during LlamaIndex ingestion.
The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.