DeepSeek · Current Flash generation
DeepSeek V4.1 Flash API pricing & cost
Peak/off-peak token pricing, cache-hit rates, context limits and direct links into TokenCOGS workload calculators.
Blended input / 1M$0.194
Blended output / 1M$0.775
Blended cache hit / 1M$0.0039
Context window1M tokens
TokenCOGS uses its existing 7 peak / 17 off-peak weekday blend for single-number calculator estimates. Exact rates are below.
Exact provider rates
Peak vs off-peak
| Rate | Off-peak | Peak |
|---|---|---|
| Input cache miss | $0.15 / 1M | $0.30 / 1M |
| Input cache hit | $0.003 / 1M | $0.006 / 1M |
| Output | $0.60 / 1M | $1.20 / 1M |
DeepSeek publishes peak hours on weekdays only; weekends are off-peak, so exact invoice cost depends on request timing.
Model facts
What changed
- Released September 10, 2026.
- API model name:
deepseek-flash. - Context window: 1M tokens; maximum output: 384K tokens.
- Legacy
deepseek-v4-flashrequests are routed to V4.1 Flash and billed at V4.1 Flash rates.
Provider note
V4 Pro is still available
DeepSeek initially announced a September 14 phase-out for V4 Pro, then reversed that plan in its September 10 changelog. TokenCOGS therefore keeps V4 Pro as a current model and tracks V4.1 Flash alongside it.
Compare DeepSeek models →Use your workload