📋 Deepseek loss leader

by anon · 2026-08-06 15:07:15
← all clips 🖨️ print / PDF
Table of contents Bottom line What “DeepSeek Flash” is—and why its cost can be low Is the model actually comparable in capability? Independent evaluation Caution on “intelligence score” What do independent hosts actually charge? Observed third-party V4 Flash prices Interpretation Can it be profitable? A transparent serving-economics check Revenue per million tokens GPU cost envelope Required throughput to cover that cost The key qualification: utilization Why the price can be lower than comparable “intelligence” models What could make it unprofitable or force a price increase? 1. Long-context economics are separate from normal-token economics 2. FP4 quality/performance and operational complexity 3. High SLA / enterprise deployment costs more 4. The low end is less proven than the mainstream rate Verdict for the argument you are settling Claim: “DeepSeek Flash can actually be hosted profitably for the price it’s at across the board.” Claim: “DeepSeek Flash pricing is an outlier versus similarly intelligent models.” Sources

🧑 You

Is deepseek flash hosting costs an outlier compared to other models of similar size and intelligence scores? Do some deep research on the topic. Ignore the cached input pricing on deeepseek's own platform, we already know this is an anomaly and they're correcting it by raising prices. Fireworks.ai pricing seems close to the rest of the industry for hosting it. The arguement we're trying to settle is if this model can actually be hosted profitably for the price it's at across the board

▶ Tool: read_file [path='/Users/patw/skills/skill_index.md']

▶ Tool: todowrite [todos={}]

▶ Tool: web_search [query='DeepSeek Flash model pricing hosting Fireworks AI De]…

▶ Tool: fetch_url [url='https://fireworks.ai/models/fireworks/deepseek-v4-flas]…

▶ Tool: fetch_url [url='https://openrouter.ai/deepseek/deepseek-v4-flash']

▶ Tool: web_search [query='DeepSeek V4 Flash technical report 284B 13B active b]…

▶ Tool: web_search [query='DeepSeek V4 Flash providers pricing Together AI Repl]…

▶ Tool: web_search [query='Fireworks AI DeepSeek V4 Flash pricing per million t]…

▶ Tool: web_search [query='DeepSeek V4 Flash benchmark comparison Qwen 3.5 GPT ]…

▶ Tool: todowrite [todos={}]

▶ Tool: fetch_url [url='https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash']

▶ Tool: fetch_url [url='https://artificialanalysis.ai/models/deepseek-v4-flash']

▶ Tool: fetch_url [url='https://deepinfra.com/blog/deepseek-v4-flash-vs-qwen3-]…

▶ Tool: web_search [query='GPU cloud H100 H200 B200 hourly rental price 2026 in]…

▶ Tool: web_search [query='DeepSeek V4 Flash inference throughput tokens per se]…

▶ Tool: web_search [query='DeepSeek V4 Flash Hugging Face model card BF16 FP8 w]…

▶ Tool: todowrite [todos={}]

▶ Tool: todowrite [todos={}]

🤖 Assistant

Bottom line

No: DeepSeek V4 Flash’s third-party hosting price does not look like a loss-leading outlier. It is an unusually strong price/performance outlier, but the evidence is consistent with it being hostable profitably at the broadly observed market price—roughly $0.09–$0.14/M input and $0.18–$0.28/M output—provided the host has competent inference engineering, high utilization, and sensible long-context/KV-cache controls.

The important distinction is:


What “DeepSeek Flash” is—and why its cost can be low

I treated the current V4 Flash family as the object of the question (preview 0423 and current 0731 are materially the same serving class).

Property

DeepSeek V4 Flash

Total parameters

284B

Active parameters/token

13B

Weights

MIT / commercial use permitted

Native context

1M tokens

Attention

Hybrid CSA + HCA sparse/compressed attention

Published release format

FP4 expert weights + FP8 other weights

Approx. serving weight footprint

~160 GB in native mixed precision / ~284 GB FP8-only

Sources: DeepSeek model card / technical report, vLLM recipe, Lambda deployment notes. The economic point is that 284B total parameters is not the correct numerator for decode compute. Per generated token, only 13B parameters are routed. That is approximately:

There is still a real fixed memory/distribution cost: all experts must be resident across a multi-GPU node. But the MoE route reduces the expensive per-token compute and memory-bandwidth work. The hybrid attention design also materially reduces long-context work and KV memory: DeepSeek reports V4 Pro needs just 27% of V3.2’s single-token inference FLOPs and 10% of its KV cache at 1M context; Flash shares the same hybrid-attention design. That does not mean Flash has precisely those ratios, but it supports the premise that this architecture is deliberately designed to reduce serving cost.


Is the model actually comparable in capability?

Yes, and this makes the price/performance result genuinely striking rather than a small-model price masquerading as frontier quality.

Independent evaluation

Artificial Analysis rates the current Flash 0731 max-reasoning version:

For perspective, in its published comparison data:

Model

AA Intelligence

Active params

Public list price, input/output per M

DeepSeek V4 Flash 0731 Max

49.9

13B

$0.14 / $0.28

MiniMax M3

44.4

23B

$0.30 / $1.20

MiMo V2.5 Pro

42.2

42B

$0.435 / $0.87

DeepSeek V4 Pro Max

44.3

49B

$0.435 / $0.87

GPT-OSS-120B high

23.8

5.1B

$0.15 / $0.60

Nemotron 3 Ultra

37.8

55B

$0.60 / $2.75

Kimi K2.6

44.2

32B

$0.95 / $4.00

That comparison includes vendor list prices, not necessarily every host’s current price, but it establishes the central fact: Flash’s normal price is 3–10× below credible peers on output tokens while its measured intelligence is at or above them. That is an outlier in customer value. It is not necessarily an outlier in infrastructure cost, because V4 Flash has by far the smallest active model among the high-intelligence open models shown above.

Caution on “intelligence score”

Reasoning effort matters a lot. Flash’s quality is materially lower with non-thinking enabled. The DeepSeek model card reports, for example:

Metric

Flash non-think

Flash high

Flash max

GPQA Diamond

71.2

87.4

88.1

LiveCodeBench

55.2

88.4

91.6

SWE-Bench Verified

73.7

78.6

79.0

So the model’s exceptional headline capability is purchased partly by emitting a lot of reasoning tokens. That shifts revenue toward the output-token side—which is also the side that most directly pays for decode cost.


What do independent hosts actually charge?

Excluding DeepSeek’s own cache anomaly, the actual external market is unusually consistent.

Observed third-party V4 Flash prices

Provider / route

Input $/M

Output $/M

Serving format / note

DeepInfra

$0.09

$0.18

FP4

Sail Research via OpenRouter

$0.09

$0.18

FP4

StreamLake / Baidu via OpenRouter

~$0.088

~$0.176

FP8

GMICloud via OpenRouter

~$0.094

~$0.188

FP8

DeepInfra, prior published direct rate

$0.10

$0.20

FP4

SiliconFlow

$0.13

$0.28

FP8

Fireworks

$0.14

$0.28

serverless

Novita

$0.14

$0.28

FP8

Parasail

$0.14

$0.28

FP8

Alibaba Cloud Intl.

$0.134

$0.268

FP8

Venice

$0.138

$0.275

Morph

$0.139

$0.278

Sources: OpenRouter model/provider page, Fireworks model page, DeepInfra comparison. OpenRouter listed 21 hosting providers for the preview model at the time observed. That does not prove each endpoint is profitable, but it is meaningful evidence against a single-provider subsidy explanation:

  1. Providers span China/Asia and US-based operators.
  2. Both FP4 and FP8 endpoints exist.
  3. The $0.14/$0.28 tier recurs across independent commercial operators.
  4. The cheaper $0.09/$0.18 tier also recurs among multiple operators, rather than appearing only at DeepSeek.

Interpretation


Can it be profitable? A transparent serving-economics check

Without a host’s GPU contracts, utilization, batch profile, and model-parallel configuration, no outside analyst can calculate exact gross margin. But we can test whether the price requires impossible throughput.

Revenue per million tokens

At the broadly used $0.14 input / $0.28 output rate:

The output rate is what matters most for high-effort reasoning workloads, because reasoning tokens are billed as output and dominate decode resource use.

GPU cost envelope

Published 2026 market references show roughly:

Sources: JarvisLabs H100 pricing survey, JarvisLabs H200 survey. Flash needs a multi-GPU configuration because of its ~160 GB native mixed-precision weights plus runtime and KV cache. An 8-GPU H100/H200-style serving group is a defensible conservative mental model for high-quality serving, though it is not necessarily the configuration every provider uses. Assume a fully loaded cluster cost of:

Required throughput to cover that cost

For output price \(P_o\), cluster cost \(C\), break-even aggregate decode throughput \(T\): \[ T = \frac{C}{3600} \times \frac{1{,}000{,}000}{P_o} \] At $0.28/M output:

Fully loaded 8-GPU group cost

Break-even output throughput

$25/hour

24.8 tok/s aggregate

$35/hour

34.7 tok/s aggregate

$50/hour

49.6 tok/s aggregate

At the aggressive $0.18/M output tier:

Fully loaded 8-GPU group cost

Break-even output throughput

$25/hour

38.6 tok/s aggregate

$35/hour

54.0 tok/s aggregate

$50/hour

77.2 tok/s aggregate

Those requirements are not absurd. They are far below the aggregate throughput that a well-batched 8-GPU modern accelerator group should be able to sustain for a 13B-active MoE, assuming normal context lengths and enough concurrent demand. A provider must preserve interactive per-request latency, so it cannot run at maximum offline batch throughput; nevertheless the break-even bar at $0.28/M is modest. A useful sanity check: Artificial Analysis measures about 102 output tok/s per streamed request on the current first-party API. That is not an aggregate hardware throughput benchmark and should not be plugged directly into a cost model, but it confirms that customers receive high per-request speed rather than a model that is obviously compute-bound at a few tokens per second.

The key qualification: utilization

The price only works if the host operates a pool, not a mostly idle dedicated pod. At $0.28/M output and $35/hour fully-loaded cost:

Thus:

The market price is viable for a multi-tenant, batched serverless pool. It is not evidence that an arbitrary buyer can deploy a dedicated low-volume V4 Flash endpoint at that price and make money.


Why the price can be lower than comparable “intelligence” models

The apparent paradox dissolves once the comparison uses serving-relevant dimensions:

  1. Active 13B MoE path.

Flash’s active path is dramatically smaller than peer frontier open models. This cuts decode compute and HBM reads per output token.

  1. FP4/FP8 native release.

The official instruct release uses FP4 experts and FP8 elsewhere. Multiple hosts explicitly serve FP4 or FP8. This lowers weight residency and memory traffic relative to BF16.

  1. Hybrid sparse/compressed attention.

Especially for long context, the architecture reduces context processing/KV pressure relative to conventional attention designs.

  1. Open weights, MIT license.

Hosts pay no per-token model royalty to DeepSeek. Their costs are infrastructure and operations, not a stacked model-provider margin.

  1. High-volume multi-tenant optimization.

Continuous batching, prefix caching, routing, expert parallelism, and KV-cache management make a huge difference. The retail API rate should be evaluated as a pool-level average, not as the price of one serialized request.

  1. Output-heavy reasoning billing.

The expensive work of generating chain-of-thought is billed as output tokens. Even if the model is verbose, that can improve the operator’s economics if output price is set appropriately.


What could make it unprofitable or force a price increase?

This is the other side of the argument: the price is plausible, but it is not invulnerable.

1. Long-context economics are separate from normal-token economics

A “1M context supported” claim does not mean a provider can profitably accept unlimited 1M-token prompts at the same flat input rate and latency target. Even with compressed attention/KV improvements:

Hosts need tiering, rate controls, context limits, caching policy, admission control, or account-level utilization controls. The listed public price is best interpreted as an average over normal production traffic, not an all-you-can-eat 1M-token commitment.

2. FP4 quality/performance and operational complexity

The cheap endpoints commonly use FP4. That is a standard commercial choice, but providers may differ in:

Equal token price does not mean perfectly equal service.

3. High SLA / enterprise deployment costs more

Fireworks’ public page distinguishes serverless from on-demand dedicated deployment. A serverless market price can be profitable while a dedicated, reserved-capacity, no-rate-limit deployment has completely different economics.

4. The low end is less proven than the mainstream rate

$0.09/$0.18 may be sustainable for some large low-cost providers, but it is more likely to change than $0.14/$0.28. The repeated $0.14/$0.28 price across Fireworks, Novita, Parasail, and others is the strongest “normal commercial” reference point.


Verdict for the argument you are settling

Claim: “DeepSeek Flash can actually be hosted profitably for the price it’s at across the board.”

Verdict: Mostly yes, with an important narrowing.

Claim: “DeepSeek Flash pricing is an outlier versus similarly intelligent models.”

Verdict: Yes in customer value; no in a way that proves third-party hosting is loss-making. It is exceptionally cheap per quality point, possibly the strongest open-weight value point in the market. But its 13B active MoE, FP4/FP8 serving format, and efficient attention explain why its infrastructure cost can be structurally lower than the much larger-active-parameter models it beats or matches on intelligence. The right conclusion is:

DeepSeek V4 Flash is a technology/price-performance outlier, not clearly a hosting-economics anomaly.
Its normal external price is believable as profitable at scale. The most aggressive rates deserve skepticism about durability, but the industry-consensus $0.14/$0.28 rate does not.

Sources

  1. DeepSeek V4 Flash official Hugging Face model card and technical report
  2. Artificial Analysis: DeepSeek V4 Flash 0731 intelligence, price, speed, peer comparisons
  3. Fireworks: V4 Flash serverless price and model details
  4. OpenRouter: provider inventory and provider-level pricing
  5. DeepInfra: V4 Flash vs Qwen3.6 and GLM-4.6 serving comparison
  6. Lambda: V4 Flash deployment notes / FP8 model format