> ## Documentation Index
> Fetch the complete documentation index at: https://docs.codeflare.cc/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat Completions

> Use messages with the OpenAI-compatible chat endpoint.

## Basic call

Use `POST /v1/chat/completions` with a `chat-completions` model, Bearer authentication, a `messages` array, and a reasonable `max_completion_tokens` limit.

<CodeGroup>
  ```bash cURL theme={null}
  curl "https://YOUR-REGION.example/v1/chat/completions" \
    -H "Authorization: Bearer $CODEFLARE_API_KEY" \
    -H 'Content-Type: application/json' \
    -d '{"model":"MODEL_ID","messages":[{"role":"user","content":"Hello"}],"max_completion_tokens":256}'
  ```

  ```python Python theme={null}
  from openai import OpenAI
  import os

  client = OpenAI(
      api_key=os.environ["CODEFLARE_API_KEY"],
      base_url="https://YOUR-REGION.example/v1", max_retries=0,
  )
  response = client.chat.completions.create(
      model="MODEL_ID",
      messages=[{"role": "user", "content": "Hello"}],
      max_completion_tokens=256,
  )
  print(response.choices[0].message.content)
  ```

  ```javascript JavaScript theme={null}
  import OpenAI from "openai";
  const client = new OpenAI({
    apiKey: process.env.CODEFLARE_API_KEY,
    baseURL: "https://YOUR-REGION.example/v1", maxRetries: 0,
  });
  const response = await client.chat.completions.create({
    model: "MODEL_ID", messages: [{ role: "user", content: "Hello" }],
    max_completion_tokens: 256,
  });
  console.log(response.choices[0].message.content);
  ```
</CodeGroup>

## Streaming

Set `stream: true`. In Python, iterate over `stream` and read nonempty `chunk.choices[0].delta.content`. In JavaScript, use `for await` and read `choices[0]?.delta?.content`. Native HTTP SSE ends with `[DONE]`; do not use a Responses `response.output_text.delta` event parser.

## Choose the right protocol

Chat Completions has different input and output formats from Responses. Codex CLI uses Responses; Claude Code uses Messages. For platform response storage, lookup, and continuation, see [Responses](/en/api/responses).

Tools, structured output, and multimodal input depend on model and upstream support. Disable automatic retries and check logs before generating again after disconnection.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.