> ## Documentation Index
> Fetch the complete documentation index at: https://docs.codeflare.cc/llms.txt
> Use this file to discover all available pages before exploring further.

# Responses

> Generate and stream output, store and continue responses, and cancel background work.

## Basic call

Use `POST /v1/responses`, Bearer authentication, and a `responses` model. This example does not store response content on the platform.

<CodeGroup>
  ```python Python theme={null}
  from openai import OpenAI
  import os

  client = OpenAI(
      api_key=os.environ["CODEFLARE_API_KEY"],
      base_url="https://YOUR-REGION.example/v1",
      max_retries=0,
  )
  response = client.responses.create(
      model="MODEL_ID", input="Hello",
      max_output_tokens=256, store=False,
  )
  print(response.output_text)
  ```

  ```javascript JavaScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: process.env.CODEFLARE_API_KEY,
    baseURL: "https://YOUR-REGION.example/v1",
    maxRetries: 0,
  });
  const response = await client.responses.create({
    model: "MODEL_ID", input: "Hello",
    max_output_tokens: 256, store: false,
  });
  console.log(response.output_text);
  ```

  ```bash cURL theme={null}
  curl "https://YOUR-REGION.example/v1/responses" \
    -H "Authorization: Bearer $CODEFLARE_API_KEY" \
    -H 'Content-Type: application/json' \
    -d '{"model":"MODEL_ID","input":"Hello","max_output_tokens":256,"store":false}'
  ```
</CodeGroup>

## Streaming

Set `stream: true` and process individual events instead of treating SSE as JSON. Reuse the configured `client`:

<CodeGroup>
  ```python Python theme={null}
  with client.responses.create(
      model="MODEL_ID", input="Hello",
      max_output_tokens=256, store=False, stream=True,
  ) as stream:
      for event in stream:
          if event.type == "response.output_text.delta":
              print(event.delta, end="", flush=True)
  ```

  ```javascript JavaScript theme={null}
  const stream = await client.responses.create({
    model: "MODEL_ID", input: "Hello",
    max_output_tokens: 256, store: false, stream: true,
  });
  for await (const event of stream) {
    if (event.type === "response.output_text.delta") {
      process.stdout.write(event.delta);
    }
  }
  ```
</CodeGroup>

## Storage and continuation

`store=true` retains content for 30 days by default. Keep the response ID. With the same user and appropriate model permissions, query `GET /v1/responses/RESPONSE_ID`, inspect `GET /v1/responses/RESPONSE_ID/input_items`, and continue with `previous_response_id` through the original region.

`store=false` keeps no platform content copy and does not permit HTTP lookup or continuation. History is available only within the same WebSocket connection. Background execution requires `store=true`.

## Background execution and cancellation

When a channel supports `background`, submit with `background: true, store: true` and poll the stored response state. Cancel a background response:

```bash theme={null}
curl -X POST "https://YOUR-REGION.example/v1/responses/RESPONSE_ID/cancel" \
  -H "Authorization: Bearer $CODEFLARE_API_KEY"
```

`DELETE /v1/responses/RESPONSE_ID` cancels execution, denies future lookup and continuation, and removes content while retaining financial records. Cancellation or disconnection does not imply zero usage.

## Other capabilities

`POST /v1/responses/compact`, `POST /v1/responses/input_tokens`, and WebSocket `GET /v1/responses` require channel support. Compact is metered; token counting has no generation charge. Event recovery uses the `stream` and `starting_after` query parameters in the original region.

<Warning>After an interruption, check the stored response and logs before sending another generation request. See [Capabilities and limits](/en/capabilities) for tool and multimodal support.</Warning>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.