1. Reveal follows queue pressure
Model output does not arrive at a stable rate. Small chunks should keep a measured cadence; a growing queue has to catch up. The reveal engine calculates a speed from the current backlog and carries fractional character debt between frames. At 60Hz, the same function produces these decisions:
| Pending characters | Target speed | Characters this frame |
|---|---|---|
| 8 | 101.4 chars/s | 1 |
| 32 | 154.7 chars/s | 2 |
| 128 | 456.0 chars/s | 7 |
| 512+ | 600.0 chars/s | 10 |
积压少时保持从容节奏,积压增加后逐步加速;速度有上限,避免一次把长段内容倾倒到页面。
2. One spring owns the vertical motion
Each layout change updates the target height. A damped spring carries its position and velocity into the next frame, so new wraps, code blocks, tables, and tool results feed one continuous trajectory. Physical time is clamped after a long main-thread stall instead of replaying the entire missed interval in one paint.
The same initial lag settles in nearly the same wall-clock time at 60Hz and 120Hz:
| Initial lag | 60Hz | 120Hz |
|---|---|---|
| 16px | 750ms | 725ms |
| 48px | 883ms | 858ms |
| 96px | 967ms | 950ms |
| 192px | 1067ms | 1033ms |
刷新率改变时帧数会变化,但实际运动时间接近,因此 60Hz 与 120Hz 下保持相似手感。
3. Reveal and follow share pressure
The follower keeps a small amount of measured room for the next wrap. When safe visual lag fills, it scales reveal pressure from 1.0 toward 0.55. This slows incoming layout growth instead of letting text outrun the scroll spring.
| Lag inside a 48px capacity | Reveal scale |
|---|---|
| 0–12px | 1.000 |
| 24px | 0.775 |
| 36–48px | 0.550 |
4. Why the follow is zero-reflow
The follower never writes top, left, width, or scrollTop-style layout properties during streaming. It compensates paint position with a single compositor transform (warmed up with will-change) on each outermost message surface, so no frame during a follow touches the layout tree — the grow/paint axis is 0-reflow, and the only repaint is the newly revealed text itself.
Per-frame visual displacement is also bounded: wrap compensation is rate-limited to a fixed step (FOLLOW_PAINT_SHIFT_MAX_STEP_PX) instead of one full line in a single paint, so a 24 px wrap lands as ≤8px per frame. This is what makes a fast stream read as continuous motion rather than a page hopping one line at a time.
The checked-in browser gates reproduce the feel end-to-end:
| Gate | Command | Passing result |
|---|---|---|
| Render audit (5 streaming scenarios) | node scripts/run-render-audit.mjs | 10/10 clean, zero displacement regressions |
| Overflow / rebound | node scripts/verify-overflow.mjs --runs 3 | 3/3 no over-scroll, no rebound |
| Tail smoothing | node scripts/probe-tailbob.mjs 1200 20 | frame step ≤30px, 7-frame amplitude ≤32px |
跟随路径只写合成层 transform,主线程布局零参与;换行补偿单帧限幅 ≤8px,高速流也零跳帧。以上闸门均可在仓库内一键复现。
Reproduce the benchmark
git clone https://github.com/Laplace-bit/dsh-smooth-stream.git
cd dsh-smooth-stream
pnpm install
pnpm benchmark
Run pnpm benchmark to regenerate the numbers on your machine (the source is stream-engine.ts). A reference recording on Apple M5, Node.js v22.22.1, macOS arm64 measured ~47.3 million queue decisions/s and ~89.3 million spring decisions/s as the median of seven runs after two warm-ups.
Scope: this microbenchmark measures pure TypeScript decisions only. It does not measure React commits, Markdown parsing, browser layout, paint, device thermals, or network time. Those require a Performance trace in a real DeepSeek Harness session. The high operation counts only show that the two math functions are not the likely UI bottleneck; the browser-level gates above cover the end-to-end feel.
What should be measured next
The useful end-to-end follow-up is a shared Chrome Performance trace for the same long Markdown fixture on desktop and a lower-power mobile device. That trace should report scripting, style/layout, paint, dropped frames, and the exact Harness/plugin versions. Until that exists, this project will not publish a broad “X% faster” claim.