TESSFeed

Technical AI signals · Daily at 8 PM ET

Tuesday, August 18, 2026

Speculative Decoding and Agent Runtimes

Today was an operator-heavy day led by a concrete speculative decoding challenge to DeepSeek’s DSpark, a serious new agent runtime result on Terminal-Bench, and a wave of shippable agent infrastructure from Vercel, NVIDIA, and LangSmith.

🛰️  Top Signals

X / Twitter

Blogs

Hacker News

Research

MOSS-VL Technical Report

MOSS-VL uses gated cross-attention for streaming vision and a staged curriculum, and it leads temporal-reasoning video sets while staying competitive offline at comparable scale.

OpenMOSS · arXiv · code · project

YouTube

Google’s AI LEGENDS Built a Startup

The video says Discovery Loop was founded by Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals to automate experiment design, testing, and evaluation.

Radio

Hugging Face

deepseek-ai/DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 adds a DSpark speculative decoding module and is described as broadly competitive with top proprietary models.

model

deepseek-ai/DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 uses the same speculative-decoding structure as DeepSeek-V4-Flash-DSpark and has a smaller activated parameter count than the preview Pro model.

model

moonshotai/Kimi-K3

Kimi K3 is a 2.8T-parameter open-weight multimodal model with a 1-million-token context window.

model

Qwen/Qwen3.8-27B

Qwen3.8-27B is compatible with Transformers, vLLM, SGLang, and TokenSpeed, and Qwen Cloud plans a hosted version with 1M context length and built-in tools.

model

Qwen/Qwen3.8-27B-FP8

Qwen3.8-27B-FP8 uses fine-grained FP8 quantization with block size 128 and claims nearly identical performance to the original model.

model

unsloth/Qwen3.8-27B-GGUF

Unsloth's Qwen3.8-27B GGUF adds developer-role support, MTP fast inference, and improved nested-object tool calling.

model

Qwen/Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is the base for Qwen3.8-Max, which adds vision input, non-thinking mode, 1M context, and built-in tools.

model

GitHub

jundot/omlx

oMLX is a Mac LLM server with continuous batching and tiered KV caching managed from the menu bar.

Python · ★ 19,371

akitaonrails/ai-memory

ai-memory lets coding agents keep long-term memory across tools, including Claude Code, OpenAI Codex, and directory handoffs.

Rust · ★ 2,676

volcengine/OpenViking

OpenViking stores agent context as a virtual filesystem under viking:// and loads content in L0, L1, and L2 tiers.

Python · ★ 29,318

chaitanyagiri/munder-difflin

Munder Difflin wraps Claude Code, Codex, Grok, Kimi Code, Qwen, OpenCode, Crush, pi.dev, and GitHub Copilot CLI into a multi-agent harness.

TypeScript · ★ 1,982

NawfalMotii79/PLFM_RADAR

AERIS-10 is an open-source 10.5 GHz phased-array radar with 3km and 20km versions.

PLSQL · ★ 24,276

bojieli/ai-agent-book

The book is organized into 10 chapters and 103 companion experiments, with versions in 14 languages.

Python · ★ 39,074

harry0703/MoneyPrinterTurbo

MoneyPrinterTurbo generates video scripts, matches素材, adds subtitles and background music, and outputs short videos from a topic or keyword.

Python · ★ 108,445

basecamp/omarchy

Omarchy is a Linux distribution with a manual covering navigation, hotkeys, clipboard history, and CLI tools.

Shell · ★ 26,390

Get this in your inbox

The same feed, delivered daily at 8 PM ET. No spam, unsubscribe anytime.