> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gregapi.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat Completions

> Use GregAPI's OpenAI-compatible chat completion and tool-calling interface.

`POST /v1/chat/completions` is the most commonly used chat generation endpoint. It supports streaming output, function calling (`tools`/`functions`), and JSON mode.

## Request example

```bash theme={null}
curl -X POST "https://api.gregapi.com/v1/chat/completions" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [
      {"role": "system", "content": "You are the GregAPI product assistant."},
      {"role": "user", "content": "Introduce GregAPI in three sentences."}
    ],
    "temperature": 0.7
  }'
```

## Common usage

* **Streaming**: set `stream: true` in the request body. Add `-N` for cURL, or use `requests.post(..., stream=True)` in Python to read line by line.
* **Function calling**: describe callable functions via `tools` and `tool_choice`, then parse `tool_calls` in the response and run your business logic.
* **JSON constraint**: set `response_format` to `{"type": "json_schema"}` to have the model return structured data strictly.

GregAPI automatically aligns differences across common compatible interfaces, reducing adaptation cost when switching models.

## GregAPI extension fields

The following fields are GregAPI's extensions on top of the standard OpenAI response body, used to expose billing details of the upstream model.

### `usage.prompt_tokens_details`

| Field | Meaning | Applicable models |
| - | - | - |
| `cached_tokens` | Number of input tokens that hit the cache | All cache-enabled models |
| `cached_write_tokens` | Total cache writes (= 5m + 1h) | Claude |
| `cached_write_5m_tokens` | Tokens written to the 5-minute cache | Claude |
| `cached_write_1h_tokens` | Tokens written to the 1-hour cache | Claude |

`prompt_tokens` is the total input tokens (including cache hits and cache writes). Cache-write fields are `omitempty`, so models without cache writes (e.g. GPT) do not output them.

```json theme={null}
{
  "usage": {
    "prompt_tokens": 58518,
    "prompt_tokens_details": {
      "cached_tokens": 11944,
      "cached_write_tokens": 46109,
      "cached_write_5m_tokens": 29762,
      "cached_write_1h_tokens": 16347
    }
  }
}
```

Combined with the unit-price fields in billing reconciliation, you can estimate the cache-write cost of a single request directly from the response body.


## OpenAPI

````yaml openapi/llm-en.yaml POST /v1/chat/completions
openapi: 3.0.3
info:
  title: GregAPI Large Language Models
  version: 1.0.0
  description: >-
    Large language model (LLM) endpoints behind the GregAPI unified gateway,
    covering general chat completion, multi-modal responses, and native
    message/generation protocols.
servers:
  - url: https://api.gregapi.com
security:
  - bearerAuth: []
paths:
  /v1/chat/completions:
    post:
      summary: Chat Completions
      description: >-
        Chat completion and tool calling through GregAPI's OpenAI-compatible
        interface, with streaming, function/tool calling, and JSON mode.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - model
                - messages
              properties:
                model:
                  type: string
                  description: Model name
                messages:
                  type: array
                  description: Conversation messages
                  items:
                    type: object
                    properties:
                      role:
                        type: string
                        enum:
                          - system
                          - user
                          - assistant
                          - tool
                      content:
                        type: string
                temperature:
                  type: number
                  description: Sampling temperature, default 1
                stream:
                  type: boolean
                  description: Whether to stream output
                tools:
                  type: array
                  description: Callable function definitions
                tool_choice:
                  type: string
                response_format:
                  type: object
                  description: Output format, e.g. JSON Schema (json_schema)
            example:
              model: gpt-4o-mini
              messages:
                - role: system
                  content: You are GregAPI's product assistant.
                - role: user
                  content: Introduce GregAPI in three sentences.
              temperature: 0.7
      responses:
        '200':
          description: Successful chat completion response
components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.