• Playground
  • Models
  • Support

Zero data retentionYour prompts and outputs are never stored or trained on

© 2026 BoundlessTermssupport@boundless.network
All models

DeepSeek V4.1 Flash

by DeepSeek ~ Fast, low-cost agentic coding and visual analysis

DeepSeek V4.1 Flash is a 552B-parameter mixture-of-experts model with a causal encoder-decoder architecture: 8B parameters activate for input tokens and 16B for output tokens. It combines fast, low-cost agentic coding and reasoning with native visual understanding, tool use, and a 1M-token context window. Its substantially smaller KV cache also lowers the cost of long-running, cache-heavy workloads.

1M context
  • Text
  • Reasoning
  • Tool use
  • Image input
  • Vision
Model weights
deepseek-v4.1-flashTry it in the playground

Pricing

Pricing
USD per 1M tokens.
WindowInputCached inputOutput
asap$0.20$0.01$1.00
What one request costs
Prompt90Ktokens
Cached prefix63Ktokens
Reply2Ktokens

$0.0075

per request to DeepSeek V4.1 Flash

Prompt
71.7%
Cached
8.4%
Reply
19.9%

DeepSeek V4.1 Flash: $0.00753 per request

90,000 prompt tokens, 63,000 of them cached, and 1,500 reply tokens.

The same request, across the lineup