Skip to content

Gemini (model family)

← All terms · AI providers and models

Also called Gemini Pro, Gemini Flash

Google's flagship family of multimodal AI models, built from the ground up to natively understand and process text, code, audio, and video simultaneously.

What it is

Unlike models that stitch separate vision and audio systems together, the Gemini family processes large amounts of multimodal data natively. They feature very large context windows (often 1M to 2M+ tokens), allowing them to ingest entire video files or massive code repositories in a single request.

When you would use it

You use Gemini when your application requires processing extreme amounts of context or needs to directly interpret video, audio, and text in the same prompt.

Common operations

  • Analyzing hour-long video files and extracting data.
  • Processing massive, cross-repository coding tasks via Google Cloud.

Related terms

Where this is taught

No learning path uses this term yet. Browse the Learning Atlas for guided sequences through related ideas.

Going deeper