Decision under test
Where should long-context KV cache live?
Compare HBM-only placement, a page-aware remote tier, and near-memory partial-state attention under one synthetic 7B GQA workload. Every result comes from versioned inputs.
- Context
- 8,192 tokens
- Batch
- 16 sequences
- Weights
- FP16
- KV cache
- FP16 GQA
Policy comparison
Four scenarios, distinct service paths
Deterministic sensitivity
Break-even and counterexample, not a universal win
| Peak link bandwidth | Break-even peak compute | Status |
|---|
Separate PyTorch MPS aggregate
Decode attention: equation check, not system validation
Apple M4/MPS median and p95 summaries only; no raw iterations. This is not CUDA, HBM/CXL, near-memory/PIM, or end-to-end validation and does not calibrate the synthetic profile.
| Validation context | Measured | Independent roofline | Attention-calibrated | Calibrated error |
|---|
Supporting calibration
Device-copy transfer equation
| Validation size | Measured | Bandwidth-only | Affine | Affine error |
|---|
Reproducible evidence
Inputs, assumptions, and failure states stay visible
| Policy | Feasible | Mean decode | Throughput | Remote-media read | Interconnect read | Page read amp. | Bottleneck |
|---|