DeepSeek-V4.1-Flash
Causal encoder–decoder · CSA2 · DeepSeekMoE
A 20-layer causal encoder supplies hidden states to a 20-layer decoder. Cache reuse and sparse indexing reduce repeated work.
- Text + native vision
- DeepSeek-ViT features and text embeddings enter Single-Pass mHC. Engram is conditional token memory, not document retrieval.
- Causal encoder–decoder
- The encoder’s final hidden states supply the decoder’s global KV projections. Each half contains 20 layers.
- Sliding-window attention
- The first two encoder layers use a local window. Deeper CSA2 layers also combine selected main KV with layer-local SWA KV.
- Hierarchical sparse indexer
- Full scores all causal positions. Selected blocks form a shared candidate pool. Reindex searches only that pool; Reuse retains the last selected indices.
- 384 routed · 6 active · 1 shared
- A router selects six expert paths per token, while a shared expert contributes alongside them. Weighted outputs recombine.
- DSpark
- A separately trained semi-autoregressive drafter and confidence-scheduled verification accelerate decoding; this is not an MTP training branch.
Production implication. Cache generation, indexing and expert execution are different costs; design the serving system around each.
Public source ↗







