umans/status/umans-deepseek-v4-flash-0731
Live · refreshes every 30s
← all models
Umans DeepSeek V4 Flash Experimental
umans-deepseek-v4-flash-0731 · DeepSeek-V4-Flash · DeepSeek
In testing
264.7tok/s
throughput · p50 · last 5 min
4.74s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent model. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-glm-5.2.

90 days agoin production since Aug 6, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
peak 317.7 tok/s · Aug 1now 315.2 tok/s
90 days agopre-release before Aug 6, 2026today
TTFT p50 · time to first token, lower is better
best 1.50s · Aug 1now 4.57s
90 days agopre-release before Aug 6, 2026today
Changelog

Events for Umans DeepSeek V4 Flash

incl. gateway-wide announcements
Aug 62026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup at exactly DeepSeek's own API pricing - $0.14 / $0.28 / $0.0028 per 1M (input / output / cache read) - the cheapest way we serve real agentic work: a 284B MoE (13B active) with a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Served on our own GPU infrastructure with high availability.