OCR 分數之外:文件 parser 真正失敗的地方
2026 年 10 月 4 日
用 1,250 頁沒有調校過的 held-out 頁面,把十二個 Reader 各自用作者自己的協定跑一遍:MinerU2.5-Pro + dots.mocr 的 composite、OCR 專用模型、open-weights VLM、兩個雲端 frontier 模型,以及每台 Mac 內建的免費 OCR。重點不是誰的平均分數漂亮,而是你的系統會在哪些頁面上直接失去資料。
OCR 分數之外:文件 parser 真正失敗的地方
2026 年 10 月 4 日
用 1,250 頁沒有調校過的 held-out 頁面,把十二個 Reader 各自用作者自己的協定跑一遍:MinerU2.5-Pro + dots.mocr 的 composite、OCR 專用模型、open-weights VLM、兩個雲端 frontier 模型,以及每台 Mac 內建的免費 OCR。重點不是誰的平均分數漂亮,而是你的系統會在哪些頁面上直接失去資料。
The Cluster You Do Not Watch
2026 年 8 月 19 日
I almost never open Grafana any more. This is the operations architecture that made that true, on under US$10 a month of infrastructure — an agent that reads everything and writes nothing directly, alerts precise enough to act on, repairs that must prove they are safe before running, and one rule that turns out to be load-bearing: a check that cannot fail is not a check. None of it depends on which Kubernetes you run.
The Shapes of Agent Memory – Files, Stores, and Experience
2026 年 8 月 12 日
An agent that remembers across sessions can keep its memory as curated markdown files, as an auto-mined structured store, or as trained experience. I measured all of them: files against a structured store under one fixed model, a store-only head-to-head across the structured lineages, and an experience bank on the agentic benchmarks where the state of the art trains memory into the weights.
Why Steering Works – The Theory Beneath X Engineering for AI Agents
2026 年 7 月 26 日
An agent run is stochastic gradient descent over solution space, and steering is how the true objective enters the loop. An optimization view and a Bayesian view that ground the industry ladder — prompt, context, harness, loop engineering — in foundation theory, and explain when a big model steering a small one beats distillation, and when it does not.
優化 Transformer 模型服務參數 – Apple Silicon GPU 的案例
2026 年 7 月 11 日
一台 MacBook Pro M4 Max、一個開源 MoE 模型,針對推論優化的一次實戰:為什麼讓 LLM 吞吐量崩跌 100 倍的是 prefill(而不是 decode),以及兩個 prefill 側的關鍵優化。
Collection Autofill 的規模化架構
2026 年 5 月 28 日
這篇文章談的是 Collection Autofill 的架構設計:一個把非結構化資料整理成 structured context 的產品介面,如何在幾百列到上千萬個 AI-computed cells 的規模下,仍然保持快速、公平且可觀測。