Skip to main content
POST
Chat Completions
POST /v1/chat/completions is the most commonly used chat generation endpoint. It supports streaming output, function calling (tools/functions), and JSON mode.

Request example

Common usage

  • Streaming: set stream: true in the request body. Add -N for cURL, or use requests.post(..., stream=True) in Python to read line by line.
  • Function calling: describe callable functions via tools and tool_choice, then parse tool_calls in the response and run your business logic.
  • JSON constraint: set response_format to {"type": "json_schema"} to have the model return structured data strictly.
GregAPI automatically aligns differences across common compatible interfaces, reducing adaptation cost when switching models.

GregAPI extension fields

The following fields are GregAPI’s extensions on top of the standard OpenAI response body, used to expose billing details of the upstream model.

usage.prompt_tokens_details

prompt_tokens is the total input tokens (including cache hits and cache writes). Cache-write fields are omitempty, so models without cache writes (e.g. GPT) do not output them.
Combined with the unit-price fields in billing reconciliation, you can estimate the cache-write cost of a single request directly from the response body.

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
model
string
required

Model name

messages
object[]
required

Conversation messages

temperature
number

Sampling temperature, default 1

stream
boolean

Whether to stream output

tools
array

Callable function definitions

tool_choice
string
response_format
object

Output format, e.g. JSON Schema (json_schema)

Response

200

Successful chat completion response