🧑 You
Is deepseek flash hosting costs an outlier compared to other models of similar size and intelligence scores? Do some deep research on the topic. Ignore the cached input pricing on deeepseek's own platform, we already know this is an anomaly and they're correcting it by raising prices. Fireworks.ai pricing seems close to the rest of the industry for hosting it. The arguement we're trying to settle is if this model can actually be hosted profitably for the price it's at across the board
▶ Tool: read_file [path='/Users/patw/skills/skill_index.md']
▶ Tool: web_search [query='DeepSeek Flash model pricing hosting Fireworks AI De]…
▶ Tool: fetch_url [url='https://fireworks.ai/models/fireworks/deepseek-v4-flas]…
▶ Tool: fetch_url [url='https://openrouter.ai/deepseek/deepseek-v4-flash']
▶ Tool: web_search [query='DeepSeek V4 Flash technical report 284B 13B active b]…
▶ Tool: web_search [query='DeepSeek V4 Flash providers pricing Together AI Repl]…
▶ Tool: web_search [query='Fireworks AI DeepSeek V4 Flash pricing per million t]…
▶ Tool: web_search [query='DeepSeek V4 Flash benchmark comparison Qwen 3.5 GPT ]…
▶ Tool: fetch_url [url='https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash']
▶ Tool: fetch_url [url='https://artificialanalysis.ai/models/deepseek-v4-flash']
▶ Tool: fetch_url [url='https://deepinfra.com/blog/deepseek-v4-flash-vs-qwen3-]…
▶ Tool: web_search [query='GPU cloud H100 H200 B200 hourly rental price 2026 in]…
▶ Tool: web_search [query='DeepSeek V4 Flash inference throughput tokens per se]…
▶ Tool: web_search [query='DeepSeek V4 Flash Hugging Face model card BF16 FP8 w]…
🤖 Assistant
No: DeepSeek V4 Flash’s third-party hosting price does not look like a loss-leading outlier. It is an unusually strong price/performance outlier, but the evidence is consistent with it being hostable profitably at the broadly observed market price—roughly $0.09–$0.14/M input and $0.18–$0.28/M output—provided the host has competent inference engineering, high utilization, and sensible long-context/KV-cache controls.
The important distinction is:
I treated the current V4 Flash family as the object of the question (preview 0423 and current 0731 are materially the same serving class).
|
Property |
DeepSeek V4 Flash |
|
Total parameters |
284B |
|
Active parameters/token |
13B |
|
Weights |
MIT / commercial use permitted |
|
Native context |
1M tokens |
|
Attention |
Hybrid CSA + HCA sparse/compressed attention |
|
Published release format |
FP4 expert weights + FP8 other weights |
|
Approx. serving weight footprint |
~160 GB in native mixed precision / ~284 GB FP8-only |
Sources: DeepSeek model card / technical report, vLLM recipe, Lambda deployment notes. The economic point is that 284B total parameters is not the correct numerator for decode compute. Per generated token, only 13B parameters are routed. That is approximately:
There is still a real fixed memory/distribution cost: all experts must be resident across a multi-GPU node. But the MoE route reduces the expensive per-token compute and memory-bandwidth work. The hybrid attention design also materially reduces long-context work and KV memory: DeepSeek reports V4 Pro needs just 27% of V3.2’s single-token inference FLOPs and 10% of its KV cache at 1M context; Flash shares the same hybrid-attention design. That does not mean Flash has precisely those ratios, but it supports the premise that this architecture is deliberately designed to reduce serving cost.
Yes, and this makes the price/performance result genuinely striking rather than a small-model price masquerading as frontier quality.
Artificial Analysis rates the current Flash 0731 max-reasoning version:
For perspective, in its published comparison data:
|
Model |
AA Intelligence |
Active params |
Public list price, input/output per M |
|
DeepSeek V4 Flash 0731 Max |
49.9 |
13B |
$0.14 / $0.28 |
|
MiniMax M3 |
44.4 |
23B |
$0.30 / $1.20 |
|
MiMo V2.5 Pro |
42.2 |
42B |
$0.435 / $0.87 |
|
DeepSeek V4 Pro Max |
44.3 |
49B |
$0.435 / $0.87 |
|
GPT-OSS-120B high |
23.8 |
5.1B |
$0.15 / $0.60 |
|
Nemotron 3 Ultra |
37.8 |
55B |
$0.60 / $2.75 |
|
Kimi K2.6 |
44.2 |
32B |
$0.95 / $4.00 |
That comparison includes vendor list prices, not necessarily every host’s current price, but it establishes the central fact: Flash’s normal price is 3–10× below credible peers on output tokens while its measured intelligence is at or above them. That is an outlier in customer value. It is not necessarily an outlier in infrastructure cost, because V4 Flash has by far the smallest active model among the high-intelligence open models shown above.
Reasoning effort matters a lot. Flash’s quality is materially lower with non-thinking enabled. The DeepSeek model card reports, for example:
|
Metric |
Flash non-think |
Flash high |
Flash max |
|
GPQA Diamond |
71.2 |
87.4 |
88.1 |
|
LiveCodeBench |
55.2 |
88.4 |
91.6 |
|
SWE-Bench Verified |
73.7 |
78.6 |
79.0 |
So the model’s exceptional headline capability is purchased partly by emitting a lot of reasoning tokens. That shifts revenue toward the output-token side—which is also the side that most directly pays for decode cost.
Excluding DeepSeek’s own cache anomaly, the actual external market is unusually consistent.
|
Provider / route |
Input $/M |
Output $/M |
Serving format / note |
|
DeepInfra |
$0.09 |
$0.18 |
FP4 |
|
Sail Research via OpenRouter |
$0.09 |
$0.18 |
FP4 |
|
StreamLake / Baidu via OpenRouter |
~$0.088 |
~$0.176 |
FP8 |
|
GMICloud via OpenRouter |
~$0.094 |
~$0.188 |
FP8 |
|
DeepInfra, prior published direct rate |
$0.10 |
$0.20 |
FP4 |
|
SiliconFlow |
$0.13 |
$0.28 |
FP8 |
|
Fireworks |
$0.14 |
$0.28 |
serverless |
|
Novita |
$0.14 |
$0.28 |
FP8 |
|
Parasail |
$0.14 |
$0.28 |
FP8 |
|
Alibaba Cloud Intl. |
$0.134 |
$0.268 |
FP8 |
|
Venice |
$0.138 |
$0.275 |
— |
|
Morph |
$0.139 |
$0.278 |
— |
Sources: OpenRouter model/provider page, Fireworks model page, DeepInfra comparison. OpenRouter listed 21 hosting providers for the preview model at the time observed. That does not prove each endpoint is profitable, but it is meaningful evidence against a single-provider subsidy explanation:
Without a host’s GPU contracts, utilization, batch profile, and model-parallel configuration, no outside analyst can calculate exact gross margin. But we can test whether the price requires impossible throughput.
At the broadly used $0.14 input / $0.28 output rate:
The output rate is what matters most for high-effort reasoning workloads, because reasoning tokens are billed as output and dominate decode resource use.
Published 2026 market references show roughly:
Sources: JarvisLabs H100 pricing survey, JarvisLabs H200 survey. Flash needs a multi-GPU configuration because of its ~160 GB native mixed-precision weights plus runtime and KV cache. An 8-GPU H100/H200-style serving group is a defensible conservative mental model for high-quality serving, though it is not necessarily the configuration every provider uses. Assume a fully loaded cluster cost of:
For output price \(P_o\), cluster cost \(C\), break-even aggregate decode throughput \(T\): \[ T = \frac{C}{3600} \times \frac{1{,}000{,}000}{P_o} \] At $0.28/M output:
|
Fully loaded 8-GPU group cost |
Break-even output throughput |
|
$25/hour |
24.8 tok/s aggregate |
|
$35/hour |
34.7 tok/s aggregate |
|
$50/hour |
49.6 tok/s aggregate |
At the aggressive $0.18/M output tier:
|
Fully loaded 8-GPU group cost |
Break-even output throughput |
|
$25/hour |
38.6 tok/s aggregate |
|
$35/hour |
54.0 tok/s aggregate |
|
$50/hour |
77.2 tok/s aggregate |
Those requirements are not absurd. They are far below the aggregate throughput that a well-batched 8-GPU modern accelerator group should be able to sustain for a 13B-active MoE, assuming normal context lengths and enough concurrent demand. A provider must preserve interactive per-request latency, so it cannot run at maximum offline batch throughput; nevertheless the break-even bar at $0.28/M is modest. A useful sanity check: Artificial Analysis measures about 102 output tok/s per streamed request on the current first-party API. That is not an aggregate hardware throughput benchmark and should not be plugged directly into a cost model, but it confirms that customers receive high per-request speed rather than a model that is obviously compute-bound at a few tokens per second.
The price only works if the host operates a pool, not a mostly idle dedicated pod. At $0.28/M output and $35/hour fully-loaded cost:
Thus:
The market price is viable for a multi-tenant, batched serverless pool. It is not evidence that an arbitrary buyer can deploy a dedicated low-volume V4 Flash endpoint at that price and make money.
The apparent paradox dissolves once the comparison uses serving-relevant dimensions:
Flash’s active path is dramatically smaller than peer frontier open models. This cuts decode compute and HBM reads per output token.
The official instruct release uses FP4 experts and FP8 elsewhere. Multiple hosts explicitly serve FP4 or FP8. This lowers weight residency and memory traffic relative to BF16.
Especially for long context, the architecture reduces context processing/KV pressure relative to conventional attention designs.
Hosts pay no per-token model royalty to DeepSeek. Their costs are infrastructure and operations, not a stacked model-provider margin.
Continuous batching, prefix caching, routing, expert parallelism, and KV-cache management make a huge difference. The retail API rate should be evaluated as a pool-level average, not as the price of one serialized request.
The expensive work of generating chain-of-thought is billed as output tokens. Even if the model is verbose, that can improve the operator’s economics if output price is set appropriately.
This is the other side of the argument: the price is plausible, but it is not invulnerable.
A “1M context supported” claim does not mean a provider can profitably accept unlimited 1M-token prompts at the same flat input rate and latency target. Even with compressed attention/KV improvements:
Hosts need tiering, rate controls, context limits, caching policy, admission control, or account-level utilization controls. The listed public price is best interpreted as an average over normal production traffic, not an all-you-can-eat 1M-token commitment.
The cheap endpoints commonly use FP4. That is a standard commercial choice, but providers may differ in:
Equal token price does not mean perfectly equal service.
Fireworks’ public page distinguishes serverless from on-demand dedicated deployment. A serverless market price can be profitable while a dedicated, reserved-capacity, no-rate-limit deployment has completely different economics.
$0.09/$0.18 may be sustainable for some large low-cost providers, but it is more likely to change than $0.14/$0.28. The repeated $0.14/$0.28 price across Fireworks, Novita, Parasail, and others is the strongest “normal commercial” reference point.
Verdict: Mostly yes, with an important narrowing.
Verdict: Yes in customer value; no in a way that proves third-party hosting is loss-making. It is exceptionally cheap per quality point, possibly the strongest open-weight value point in the market. But its 13B active MoE, FP4/FP8 serving format, and efficient attention explain why its infrastructure cost can be structurally lower than the much larger-active-parameter models it beats or matches on intelligence. The right conclusion is:
DeepSeek V4 Flash is a technology/price-performance outlier, not clearly a hosting-economics anomaly.
Its normal external price is believable as profitable at scale. The most aggressive rates deserve skepticism about durability, but the industry-consensus $0.14/$0.28 rate does not.