infery
← All models

DeepSeek models

18 models on Infery.ai.

DeepSeek Flash Latest

deepseek-flash-latest

This model always redirects to the latest model in the DeepSeek Flash family.

1M ctx

18.75 cr in / 1M75 cr out / 1M1.88 cr cached in / 1M
ChatStreamingVisionToolsJSON

DeepSeek Pro Latest

deepseek-pro-latest

This model always redirects to the latest model in the DeepSeek Pro family.

1M ctx

87.5 cr in / 1M370 cr out / 1M4.13 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V3 0324

deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.

164K ctx

31.25 cr in / 1M125 cr out / 1M
ChatStreamingToolsJSON

DeepSeek V3.1

deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt…

164K ctx

31.25 cr in / 1M119 cr out / 1M16.25 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V3.1 Terminus

deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by…

164K ctx

33.75 cr in / 1M125 cr out / 1M16.88 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V3.2

deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous…

128K ctx

36.4 cr in / 1M54.6 cr out / 1M0.910 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V3.2 Exp

deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future…

164K ctx

33.75 cr in / 1M51.25 cr out / 1M
ChatStreamingToolsJSON

DeepSeek V3.2 Reasoner

deepseek-reasoner

128K ctx

36.4 cr in / 1M54.6 cr out / 1M0.910 cr cached in / 1M
ChatStreamingTools

DeepSeek V4.1 Flash

deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family.

1M ctx

37.5 cr in / 1M150 cr out / 1M0.750 cr cached in / 1M
ChatStreamingVisionToolsJSON

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated…

1M ctx

57.2 cr in / 1M172 cr out / 1M0.910 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V4 Flash 0731

deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total.

1M ctx

7.5 cr in / 1M15 cr out / 1M1.5 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V4 Flash Latest

deepseek-v4-flash-latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

1M ctx

5 cr in / 1M12.5 cr out / 1M1.25 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V4 Flash Vision Exp

deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding…

1M ctx

27.5 cr in / 1M82.5 cr out / 1M0.875 cr cached in / 1M
ChatStreamingVisionToolsJSON

DeepSeek V4 Pro

deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting…

1M ctx

172 cr in / 1M515 cr out / 1M2.86 cr cached in / 1M
ChatStreamingToolsJSON

DeepSeek V4 Pro 0813

deepseek-v4-pro-0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek.

1M ctx

165 cr in / 1M495 cr out / 1M16.25 cr cached in / 1M
ChatStreamingToolsJSON

R1

deepseek-r1

DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens.

64K ctx

87.5 cr in / 1M313 cr out / 1M
ChatStreamingToolsJSON

R1 0528

deepseek-r1-0528

May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens.

164K ctx

62.5 cr in / 1M269 cr out / 1M43.75 cr cached in / 1M
ChatStreamingToolsJSON

R1 Distill Llama 70B

deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct, using outputs from DeepSeek R1.

131K ctx

100 cr in / 1M100 cr out / 1M
ChatStreamingJSON