infery
← All models

Google models

57 models on Infery.ai.

Gemini 2.5 Computer Use

gemini-2.5-computer-use

131K ctx

156 cr in / 1M1250 cr out / 1M
ChatStreamingVisionImage output

Gemini 2.5 Flash

gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and…

1M ctx

37.5 cr in / 1M313 cr out / 1M3.75 cr cached in / 1M125 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

1M ctx

13 cr in / 1M52 cr out / 1M1.3 cr cached in / 1M39 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini 2.5 Flash Lite Preview 09-2025

gemini-2.5-flash-lite-preview-09-2025

1M ctx

13 cr in / 1M52 cr out / 1M1.3 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 2.5 Flash Native Audio

gemini-2.5-flash-native-audio

131K ctx

62.5 cr in / 1M250 cr out / 1M375 cr audio in / 1M
ChatStreamingVisionImage output

Gemini 2.5 Flash STT

gemini-2.5-flash-stt

1M ctx

0.013 cr / min

Gemini 2.5 Flash TTS

gemini-2.5-flash-tts

8K ctx

62.5 cr in / 1M1250 cr out / 1M
Audio output

Gemini 2.5 Pro

gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

1M ctx

156 cr in / 1M1250 cr out / 1M15.63 cr cached in / 1M1250 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini 2.5 Pro Preview 06-05

gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

1M ctx

156 cr in / 1M1250 cr out / 1M15.63 cr cached in / 1M156 cr audio in / 1M156 cr image in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 2.5 Pro TTS

gemini-2.5-pro-tts

8K ctx

125 cr in / 1M2500 cr out / 1M
Audio output

Gemini 3.1 Flash Lite

gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.

1M ctx

32.5 cr in / 1M195 cr out / 1M3.25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.1 Flash-Lite Preview

gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases.

1M ctx

32.5 cr in / 1M195 cr out / 1M65 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini 3.1 Flash Live

gemini-3.1-flash-live-preview

131K ctx

93.75 cr in / 1M563 cr out / 1M375 cr audio in / 1M

Gemini 3.1 Flash TTS

gemini-3.1-flash-tts

8K ctx

125 cr in / 1M2500 cr out / 1M
Audio output

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic…

1M ctx

250 cr in / 1M1500 cr out / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini 3.1 Pro Preview Custom Tools

gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general…

1M ctx

250 cr in / 1M1500 cr out / 1M25 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.5 Flash

gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.

1M ctx

188 cr in / 1M1125 cr out / 1M18.75 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.

1M ctx

37.5 cr in / 1M313 cr out / 1M3.75 cr cached in / 1M37.5 cr audio in / 1M37.5 cr image in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.6 Flash

gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development.

1M ctx

93.75 cr in / 1M469 cr out / 1M9.38 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.7 Flash

gemini-3.7-flash

Gemini 3.7 Flash is Google's latest and most capable Flash model, built for complex coding, agentic workflows and reliable multi-step…

1M ctx

93.75 cr in / 1M469 cr out / 1M9.38 cr cached in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3.8 Flash

gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.

1M ctx

93.75 cr in / 1M469 cr out / 1M9.38 cr cached in / 1M93.75 cr audio in / 1M93.75 cr image in / 1M
ChatStreamingVisionPDFToolsJSON

Gemini 3 Flash Preview

gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

1M ctx

62.5 cr in / 1M375 cr out / 1M6.25 cr cached in / 1M125 cr audio in / 1M
ChatStreamingVisionPDFToolsJSONImage output

Gemini Embedding

gemini-embedding-001

2K ctx

18.75 cr in / 1M

Gemini Embedding 2

gemini-embedding-2

8K ctx

25 cr in / 1M
VisionPDFImage output

Gemini Omni Flash 1.1 Edit

gemini-omni-flash-v1.1

Gemini Omni Flash 1.1 is Google's multimodal video model.

3.75 cr / sec
Video output

Gemini Omni Flash 1.1 Reference to Video

google/gemini-omni-flash/v1.1/reference-to-video

Gemini Omni Flash 1.1 is Google's multimodal video model.

3.75 cr / sec
Video output

Gemini Robotics-ER 1.5

gemini-robotics-er

1M ctx

37.5 cr in / 1M313 cr out / 1M125 cr audio in / 1M
ChatStreamingVisionImage output

Gemini Robotics-ER 1.6

gemini-robotics-er-1.6

131K ctx

125 cr in / 1M625 cr out / 1M250 cr audio in / 1M
ChatStreamingVisionImage output

Gemini TTS

gemini-tts

Use Gemini TTS Models to convert your prompts to real audio.

125 cr in / 1M
Music output

Gemma 2 27B

gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models.

8K ctx

81.25 cr in / 1M81.25 cr out / 1M
ChatStreamingJSON

Gemma 3 12B

gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

131K ctx

6.25 cr in / 1M18.75 cr out / 1M
ChatStreamingVisionToolsJSON

Gemma 3 27B

gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

131K ctx

10 cr in / 1M56.25 cr out / 1M5 cr cached in / 1M
ChatStreamingVisionToolsJSON

Gemma 3 4B

gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

131K ctx

6.25 cr in / 1M12.5 cr out / 1M
ChatStreamingVisionJSON

Gemma 4 26B A4B

gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. 262,144 token context window.

262K ctx

11.25 cr in / 1M37.5 cr out / 1M6.25 cr cached in / 1M
ChatStreamingVisionToolsJSON

Gemma 4 31B

gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.

262K ctx

11.25 cr in / 1M42.5 cr out / 1M6.25 cr cached in / 1M
ChatStreamingVisionToolsJSON

Google Cloud TTS

google-cloud-tts

0.00200 cr / char
Audio output

Google Gemini Flash Latest

gemini-flash-latest

This model always redirects to the latest model in the Google Gemini Flash family.

1M ctx

93.75 cr in / 1M469 cr out / 1M9.38 cr cached in / 1M93.75 cr audio in / 1M93.75 cr image in / 1M
ChatStreamingVisionPDFToolsJSON

Google Gemini Pro Latest

gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

1M ctx

250 cr in / 1M1500 cr out / 1M25 cr cached in / 1M250 cr audio in / 1M250 cr image in / 1M
ChatStreamingVisionPDFToolsJSON

Google Virtual Try On

virtual-try-on

Generate realistic virtual try-on images from a person image and a clothing product image.

9.38 cr / image
Image output

Lyria3

lyria3

Lyria 3 is most recent music model from Google

5 cr / track
Music output

Lyria 3 Clip

lyria-3-clip

1M ctx

5 cr / req
Music output

Lyria 3 Pro

lyria-3-pro

1M ctx

10 cr / req
Music output

Nano Banana

nano-banana

Nano Banana (aka Gemini 2.5 Flash Image from Google), Google's latest high quality image editing model with strong prompt adherence and…

33K ctx

37.5 cr in / 1M3750 cr out / 1M
Image output

Nano Banana 2

nano-banana-2

Generate and edit images with Google's Nano Banana 2 (Gemini 3.1 Flash Image).

66K ctx

62.5 cr in / 1M7500 cr out / 1M
Image output

Nano Banana 2 Lite

gemini-3.1-flash-lite-image

66K ctx

31.25 cr in / 1M3750 cr out / 1M
VisionImage output

Nano Banana Pro

nano-banana-pro

Nano Banana Pro (or Gemini 3 Pro Image) is Google's new state-of-the-art image generation and editing model.

131K ctx

250 cr in / 1M1500 cr out / 1M
Image output

Veo 2

veo-2

Veo 2 is Google's video generation model with realistic motion, real-world physics, and up to 4K resolution. Use Veo 2 with an API.

62.5 cr / sec
Video output

Veo 3

veo-3

Includes native audio generation, improved prompt adherence, and stunning hyperrealism.

25 cr / sec
Video output

Veo 3.1

veo-3.1

Veo 3.1 is Google's latest video generation model with synchronized audio, reference image support, and enhanced prompt adherence.

480 ctx

50 cr / sec
Video output

Veo 3.1

veo3.1-reference-to-video

Generate Videos from images using Google's Veo 3.1

50 cr / sec
Video output

Veo 3.1

veo3.1-first-last-frame-to-video

Generate videos from a first and last framed using Google's Veo 3.1

50 cr / sec
Video output

Veo 3.1 Fast

veo3.1-fast-reference-to-video

Generate videos from reference images using Google's Veo 3.1 Fast

18.75 cr / sec
Video output

Veo 3.1 Fast

veo-3.1-fast

Veo 3.1 Fast is a faster version of Google's Veo 3.1 video model with synchronized audio and high-fidelity output.

480 ctx

12.5 cr / sec
Video output

Veo 3.1 Fast

veo3.1-fast-first-last-frame-to-video

Generate videos from a first/last frame using Google's Veo 3.1 Fast

18.75 cr / sec
Video output

Veo 3.1 Lite

veo-3.1-lite

480 ctx

6.25 cr / sec
Video output

Veo3.1 Lite FLF

veo3.1-lite-first-last-frame-to-video

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

6.25 cr / sec
Video output

Veo 3 Fast

veo-3-fast

Google's Veo 3 Fast video model — the advanced AI video generation model designed for ultra-high-speed, cinematic-quality video creation.

12.5 cr / sec
Video output