Skip to main content
POST
Gemini Native Format

Endpoints

The native interface supports Authorization: Bearer <TOKEN>, and also x-goog-api-key: <TOKEN>.

Quick start

Streaming output

Multimodal input

The native interface carries multimodal content through contents[].parts[]: text uses text, and media uses inline_data (mime_type + Base64 data).

Audio understanding

Images similarly use inline_data with a mime_type of image/png, image/jpeg, and so on. The gemini-3-pro-preview-file model can analyze video files directly via URL; see the video analysis documentation for details.

Reasoning

The Gemini 2.5 and 3 series support reasoning, configured through generationConfig.thinkingConfig. The Gemini 3 series uses thinkingLevel:
The Gemini 2.5 series uses thinkingBudget:

Generation parameters

Gemini 3 recommends keeping temperature at 1.0; too low a value may degrade reasoning.

SDK example (Python · Google SDK)

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Path Parameters

model
string
required

Model name

Example:

"gemini-2.5-flash"

Body

application/json
contents
object[]
required

Conversation content

generationConfig
object

Generation configuration, e.g. temperature, thinkingConfig

Response

200

Successful generation response