Build a Billing Reconciliation Harness for Chinese AI Model Routers
python
dev.to
Build a Billing Reconciliation Harness for Chinese AI Model Routers Chinese AI APIs are moving fast enough that a static cost spreadsheet is now a liability. DeepSeek has versioned V4 Flash and V4 Pro pricing, Kimi K3 exposes separate cache-hit and cache-miss input rates for 1M-token contexts, Z.AI publishes GLM cached-input rates and tool fees, and QwenCloud uses context tiers for long prompts. If your application routes requests across these models, the invoice risk is not only "whi