infery
← All models

OpenAI models

95 models on Infery.ai.

Babbage 002

babbage-002

16K ctx

50 cr in / 1M50 cr out / 1M

ChatGPT Image (latest)

chatgpt-image-latest

625 cr in / 1M4000 cr out / 1M1000 cr image in / 1M
Image output

Codex Mini Latest

codex-mini-latest

200K ctx

188 cr in / 1M750 cr out / 1M
ChatStreamingTools

Computer Use Preview

computer-use-preview

Specialized model for computer use tool

200K ctx

375 cr in / 1M1500 cr out / 1M
ChatStreamingVision

DALL-E 2

dall-e-2

2.5 cr / image
Image output

DALL-E 3

dall-e-3

5 cr / image
Image output

Davinci 002

davinci-002

16K ctx

250 cr in / 1M250 cr out / 1M

GPT-3.5 Turbo

gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. 16,385 token context window, maximum output of 4,096 tokens.

16K ctx

62.5 cr in / 1M188 cr out / 1M
ChatStreamingToolsJSON

GPT-3.5 Turbo 16k

gpt-3.5-turbo-16k

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request…

16K ctx

375 cr in / 1M500 cr out / 1M
ChatStreamingToolsJSON

GPT-3.5 Turbo Instruct

gpt-3.5-turbo-instruct

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations.

4K ctx

188 cr in / 1M250 cr out / 1M
ChatStreamingJSON

GPT-3.5 Turbo (older v0613)

gpt-3.5-turbo-0613

GPT-3.5 Turbo is OpenAI's fastest model. 4,095 token context window, maximum output of 4,096 tokens.

4K ctx

125 cr in / 1M250 cr out / 1M
ChatStreamingToolsJSON

GPT-4

gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than…

8K ctx

3750 cr in / 1M7500 cr out / 1M
ChatStreamingToolsJSON

GPT-4.1

gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context…

200K ctx

250 cr in / 1M1000 cr out / 1M62.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-4.1 Mini

gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.

200K ctx

50 cr in / 1M200 cr out / 1M12.5 cr cached in / 1M
ChatStreamingVisionToolsJSON

GPT-4.1 Nano

gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series.

200K ctx

13 cr in / 1M52 cr out / 1M3.25 cr cached in / 1M
ChatStreamingToolsJSON

GPT-4o

gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.

128K ctx

313 cr in / 1M1250 cr out / 1M156 cr cached in / 1M12500 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-4o (2024-05-13)

gpt-4o-2024-05-13

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.

128K ctx

625 cr in / 1M1875 cr out / 1M
ChatStreamingVisionTools

GPT-4o (2024-08-06)

gpt-4o-2024-08-06

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the…

128K ctx

313 cr in / 1M1250 cr out / 1M156 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-4o (2024-11-20)

gpt-4o-2024-11-20

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve…

128K ctx

313 cr in / 1M1250 cr out / 1M156 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-4o Audio Preview

gpt-4o-audio-preview

300 cr in / 1M1200 cr out / 1M4800 cr audio in / 1M9600 cr audio out / 1M
ChatStreamingToolsJSON

GPT-4o Mini

gpt-4o-mini

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.

128K ctx

19.5 cr in / 1M78 cr out / 1M9.75 cr cached in / 1M
ChatStreamingToolsJSON

GPT-4o-mini (2024-07-18)

gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.

128K ctx

18.75 cr in / 1M75 cr out / 1M9.38 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-4o Mini Audio Preview

gpt-4o-mini-audio-preview

18 cr in / 1M72 cr out / 1M1200 cr audio in / 1M2400 cr audio out / 1M
ChatStreamingToolsJSON

GPT-4o Mini Search Preview

gpt-4o-mini-search-preview

GPT-4o mini Search Preview is a specialized model for web search in Chat Completions.

128K ctx

19.5 cr in / 1M78 cr out / 1M
ChatStreaming

GPT-4o Mini Transcribe

gpt-4o-mini-transcribe

Speech-to-text model powered by GPT-4o Mini

156 cr in / 1M625 cr out / 1M375 cr audio in / 1M

GPT-4o Mini TTS

gpt-4o-mini-tts

313 cr in / 1M1250 cr out / 1M
Audio output

GPT-4o Search Preview

gpt-4o-search-preview

GPT-4o Search Previewis a specialized model for web search in Chat Completions.

128K ctx

313 cr in / 1M1250 cr out / 1M
ChatStreaming

GPT-4o Transcribe

gpt-4o-transcribe

Speech-to-text model powered by GPT-4o

313 cr in / 1M1250 cr out / 1M750 cr audio in / 1M

GPT-4o Transcribe + Diarization

gpt-4o-transcribe-diarize

313 cr in / 1M1250 cr out / 1M750 cr audio in / 1M

GPT-4 Turbo

gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. 128,000 token context window, maximum output of 4,096 tokens.

128K ctx

1250 cr in / 1M3750 cr out / 1M
ChatStreamingVisionTools

GPT-5

gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.

200K ctx

156 cr in / 1M1250 cr out / 1M15.63 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.1

gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction…

200K ctx

156 cr in / 1M1250 cr out / 1M15.63 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.1 Codex

gpt-5.1-codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows.

200K ctx

156 cr in / 1M1250 cr out / 1M
ChatStreamingTools

GPT-5.1 Codex Max

gpt-5.1-codex-max

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.

200K ctx

156 cr in / 1M1250 cr out / 1M
ChatStreaming

GPT-5.1 Codex Mini

gpt-5.1-codex-mini

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5. 400,000 token context window, maximum output of 100,000 tokens.

200K ctx

32.5 cr in / 1M260 cr out / 1M
ChatStreaming

GPT-5.2

gpt-5.2

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.

200K ctx

219 cr in / 1M1750 cr out / 1M21.88 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.2 Chat

gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general…

128K ctx

219 cr in / 1M1750 cr out / 1M21.88 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-5.2 Codex

gpt-5.2-codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows.

200K ctx

219 cr in / 1M1750 cr out / 1M
ChatStreamingTools

GPT-5.2 Pro

gpt-5.2-pro

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro.

200K ctx

2625 cr in / 1M21000 cr out / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.3 Codex

gpt-5.3-codex

The most capable agentic coding model to date.

200K ctx

219 cr in / 1M1750 cr out / 1M21.88 cr cached in / 1M
ChatStreamingTools

GPT-5.4

gpt-5.4

A more affordable model for coding and professional work.

200K ctx

313 cr in / 1M1875 cr out / 1M31.25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.4 Mini

gpt-5.4-mini

OpenAI's strongest mini model yet for coding, computer use, and subagents

200K ctx

93.75 cr in / 1M563 cr out / 1M9.38 cr cached in / 1M
ChatStreamingVisionToolsJSONImage output

GPT-5.4 Nano

gpt-5.4-nano

OpenAI's cheapest GPT-5.4-class model for simple high-volume tasks

200K ctx

26 cr in / 1M163 cr out / 1M2.6 cr cached in / 1M
ChatStreamingToolsJSON

GPT-5.4 Pro

gpt-5.4-pro

Version of GPT-5.4 that produces smarter and more precise responses.

200K ctx

3750 cr in / 1M22500 cr out / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.5

gpt-5.5

A new class of intelligence for coding and professional work.

272K ctx

625 cr in / 1M3750 cr out / 1M62.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.5 Pro

gpt-5.5-pro

Version of GPT-5.5 that produces smarter and more precise responses.

272K ctx

3750 cr in / 1M22500 cr out / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.6 Luna

gpt-5.6-luna

GPT-5.6 model optimized for cost-sensitive workloads

272K ctx

25 cr in / 1M150 cr out / 1M2.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.6 Luna Pro

gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on…

1.1M ctx

25 cr in / 1M150 cr out / 1M2.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-5.6 Sol

gpt-5.6-sol

Flagship model for complex professional work

272K ctx

500 cr in / 1M2500 cr out / 1M50 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.6 Sol Pro

gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex…

1.1M ctx

250 cr in / 1M1250 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-5.6 Terra

gpt-5.6-terra

GPT-5.6 model that balances intelligence and cost

272K ctx

250 cr in / 1M1500 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5.6 Terra Pro

gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on…

1.1M ctx

250 cr in / 1M1500 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-5 Codex

gpt-5-codex

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows.

200K ctx

156 cr in / 1M1250 cr out / 1M
ChatStreamingTools

GPT-5 Mini

gpt-5-mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks.

200K ctx

32.5 cr in / 1M260 cr out / 1M3.25 cr cached in / 1M
ChatStreamingToolsJSON

GPT-5 Nano

gpt-5-nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low…

200K ctx

6.5 cr in / 1M52 cr out / 1M0.650 cr cached in / 1M
ChatStreamingToolsJSON

GPT-5 Pro

gpt-5-pro

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.

200K ctx

1875 cr in / 1M15000 cr out / 1M
ChatStreamingVisionPDFToolsJSONImage output

GPT-5 Search API

gpt-5-search-api

200K ctx

156 cr in / 1M1250 cr out / 1M15.63 cr cached in / 1M
ChatStreaming

GPT-6 Astra

gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. 1,050,000 token context window, maximum output of 128,000 tokens.

1.1M ctx

1250 cr in / 1M6250 cr out / 1M125 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT-6 Astra Pro

gpt-6-astra-pro

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex…

1.1M ctx

1250 cr in / 1M6250 cr out / 1M125 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT Audio

gpt-audio

300 cr in / 1M1200 cr out / 1M3840 cr audio in / 1M7680 cr audio out / 1M
ChatStreamingToolsJSON

GPT Audio 1.5

gpt-audio-1.5

300 cr in / 1M1200 cr out / 1M3840 cr audio in / 1M7680 cr audio out / 1M
ChatStreamingToolsJSON

GPT Audio Mini

gpt-audio-mini

72 cr in / 1M288 cr out / 1M1200 cr audio in / 1M2400 cr audio out / 1M
ChatStreamingToolsJSON

GPT Chat Latest

gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT.

400K ctx

625 cr in / 1M3750 cr out / 1M62.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

GPT Image 1.5

gpt-image-1.5

OpenAI's previous image generation model

625 cr in / 1M1250 cr out / 1M1000 cr image in / 1M
Image output

GPT Image 1 Mini

gpt-image-1-mini

A cost-efficient version of GPT Image 1

250 cr in / 1M1000 cr out / 1M313 cr image in / 1M
Image output

GPT Image 2

gpt-image-2

State-of-the-art image generation model

625 cr in / 1M1250 cr out / 1M1000 cr image in / 1M
Image output

gpt-oss-120b

gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic…

131K ctx

4.63 cr in / 1M21.25 cr out / 1M
ChatStreamingToolsJSON

gpt-oss-20b

gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. 131,072 token context window.

131K ctx

3.75 cr in / 1M16.25 cr out / 1M3.75 cr cached in / 1M
ChatStreamingToolsJSON

gpt-oss-safeguard-20b

gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b.

131K ctx

9.38 cr in / 1M37.5 cr out / 1M4.69 cr cached in / 1M
ChatStreamingToolsJSON

o1

o1

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding.

200K ctx

1875 cr in / 1M7500 cr out / 1M938 cr cached in / 1M
ChatStreaming

o1-mini

o1-mini

200K ctx

138 cr in / 1M550 cr out / 1M
ChatStreaming

o1-pro

o1-pro

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning.

200K ctx

18750 cr in / 1M75000 cr out / 1M
ChatStreaming

o3

o3

o3 is a well-rounded and powerful model across domains. 200,000 token context window, maximum output of 100,000 tokens.

200K ctx

250 cr in / 1M1000 cr out / 1M62.5 cr cached in / 1M
ChatStreamingTools

o3 Deep Research

o3-deep-research

OpenAI's most powerful deep research model

200K ctx

1250 cr in / 1M5000 cr out / 1M
ChatStreaming

o3-mini

o3-mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and…

200K ctx

138 cr in / 1M550 cr out / 1M68.75 cr cached in / 1M
ChatStreamingTools

o3 Mini High

o3-mini-high

OpenAI o3-mini-high is the same model as o3-mini with reasoning_effort set to high.

200K ctx

138 cr in / 1M550 cr out / 1M68.75 cr cached in / 1M
ChatStreamingPDFToolsJSON

o3-pro

o3-pro

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning.

200K ctx

2500 cr in / 1M10000 cr out / 1M
ChatStreamingTools

o4-mini

o4-mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong…

200K ctx

138 cr in / 1M550 cr out / 1M34.38 cr cached in / 1M
ChatStreamingTools

o4-mini Deep Research

o4-mini-deep-research

Faster, more affordable deep research model

200K ctx

250 cr in / 1M1000 cr out / 1M
ChatStreaming

o4 Mini High

o4-mini-high

OpenAI o4-mini-high is the same model as o4-mini with reasoning_effort set to high.

200K ctx

138 cr in / 1M550 cr out / 1M34.38 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Omni Moderation

omni-moderation-latest

33K ctx

Vision

OpenAI GPT Astra Latest

gpt-astra-latest

This model always redirects to the latest model in the OpenAI GPT Astra family.

1.1M ctx

1250 cr in / 1M6250 cr out / 1M125 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

OpenAI GPT Luna Latest

gpt-luna-latest

This model always redirects to the latest model in the OpenAI GPT Luna family.

1.1M ctx

25 cr in / 1M150 cr out / 1M2.5 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

OpenAI GPT Mini Latest

gpt-mini-latest

This model always redirects to the latest model in the OpenAI GPT Mini family.

400K ctx

93.75 cr in / 1M563 cr out / 1M9.38 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

OpenAI GPT Sol Latest

gpt-sol-latest

This model always redirects to the latest model in the OpenAI GPT Sol family.

1.1M ctx

250 cr in / 1M1250 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

OpenAI GPT Terra Latest

gpt-terra-latest

This model always redirects to the latest model in the OpenAI GPT Terra family.

1.1M ctx

250 cr in / 1M1500 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Sora 2

sora-2

Flagship video generation with synced audio

12.5 cr / sec
Video output

Sora 2 Pro

sora-2-pro

Most advanced synced-audio video generation

37.5 cr / sec
Video output

Text Embedding 3 Large

text-embedding-3-large

8K ctx

16.25 cr in / 1M

Text Embedding 3 Small

text-embedding-3-small

8K ctx

2.5 cr in / 1M

Text Embedding Ada 002

text-embedding-ada-002

8K ctx

12.5 cr in / 1M

Text Moderation

text-moderation-latest

33K ctx

TTS-1

tts-1

3750 cr in / 1M
Audio output

TTS-1 HD

tts-1-hd

0.00375 cr / char
Audio output

Whisper-1

whisper-1

0.750 cr / min