> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.predictionguard.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.predictionguard.com/_mcp/server.

# Completions

POST https://{your-pg.api-domain}.com/completions
Content-Type: application/json

Retrieve text completions based on the provided input.

Reference: https://docs.predictionguard.com/api-reference/api-reference/completions

## Authentication

- `Authorization` header (bearer token, required) — Bearer authentication of the form `Bearer <token>`, where token is your auth token.

## Request

### Body (application/json)

This endpoint expects an object.

- `model` (string, required) — The chat model to use for generating completions.
- `prompt` (CompletionsPostRequestBodyContentApplicationJsonSchemaPrompt, required)
- `frequency_penalty` (double, optional) — A value between -2.0 and 2.0, with positive values increasingly penalizing new tokens based on their frequency so far in order to decrease further occurrences.
- `logit_bias` (CompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias, optional) — Modifies the likelihood of specified tokens appearing in a response.
- `max_tokens` (integer, optional) — The maximum number of tokens in the generated completion.
- `presence_penalty` (double, optional) — A value between -2.0 and 2.0, with positive values causing a flat reduction of new tokens based on their existing presence so far in order to decrease further occurrences.
- `stop` (CompletionsPostRequestBodyContentApplicationJsonSchemaStop, optional)
- `stream` (boolean, optional) — Whether to stream back the model response.
- `stream_options` (CompletionsPostRequestBodyContentApplicationJsonSchemaStreamOptions, optional) — Extra parameters used when streaming the response.
- `temperature` (double, optional) — The temperature parameter for controlling randomness in completions. Supports a range of 0.0-2.0.
- `top_p` (double, optional) — The diversity of the generated text based on nucleus sampling. Supports a range of 0.0-1.0.
- `top_k` (integer, optional) — The diversity of the generated text based on top-k sampling.
- `output` (CompletionsPostRequestBodyContentApplicationJsonSchemaOutput, optional) — Options to affect the output of the response.
- `input` (CompletionsPostRequestBodyContentApplicationJsonSchemaInput, optional) — Options to affect the input of the request.

## Response

### 200

Successful response.

- `id` (string, optional) — Unique ID for the completion.
- `object` (string, optional) — Type of object (completion).
- `created` (integer, optional) — Timestamp of when the completion was created.
- `model` (string, optional) — The model used for generating the result.
- `choices` (list of CompletionsPostResponsesContentApplicationJsonSchemaChoicesItems, optional) — The set of result choices.

## Errors

### 400 Bad Request Error

General error response.

- `error` (string, optional) — Description of the error.

### 403 Forbidden Error

Failed auth response.

- `error` (string, optional) — Description of the error.

## Types

### CompletionsPostRequestBodyContentApplicationJsonSchemaPrompt

### CompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias

Modifies the likelihood of specified tokens appearing in a response.

- `token` (string, optional) — A string of the chosen token ID. Value is an int from -100 to 100.

### CompletionsPostRequestBodyContentApplicationJsonSchemaStop

### CompletionsPostRequestBodyContentApplicationJsonSchemaStreamOptions

Extra parameters used when streaming the response.

- `include_usage` (boolean, optional) — Whether to include tokens used in the stream response objects.

### CompletionsPostRequestBodyContentApplicationJsonSchemaOutput

Options to affect the output of the response.

- `toxicity` (boolean, optional) — Set to true to turn on toxicity processing.

### CompletionsPostRequestBodyContentApplicationJsonSchemaInput

Options to affect the input of the request.

- `block_prompt_injection` (boolean, optional) — Set to true to detect prompt injection attacks.
- `pii` (string, optional) — Set to either 'block' or 'replace'.
- `pii_replace_method` (string, optional) — Set to either 'random', 'fake', 'category', 'mask'.
- `entity_list` (list of CompletionsPostRequestBodyContentApplicationJsonSchemaInputEntityListItems, optional) — An array of entity types that the PII check should ignore.

### CompletionsPostResponsesContentApplicationJsonSchemaChoicesItems

- `index` (integer, optional) — The index position in the collection.
- `text` (string, optional) — The generated text.

### CompletionsPostRequestBodyContentApplicationJsonSchemaInputEntityListItems

## Examples

### Basic

**Request**

```json
{
  "model": "{{TEXT_MODEL}}",
  "prompt": "Will I lose my hair?",
  "frequency_penalty": 0.1,
  "logit_bias": {
    "128000": 10
  },
  "max_tokens": 1000,
  "presence_penalty": 0.1,
  "stop": "hello",
  "temperature": 1,
  "top_p": 1,
  "top_k": 50,
  "output": {
    "toxicity": true
  },
  "input": {
    "pii": "replace",
    "pii_replace_method": "random"
  }
}
```

**Response**

```json
{
  "id": "cmpl-d20e5639-ed8d-48b2-a403-08fe0fdc7add",
  "object": "text_completion",
  "created": 1727889988,
  "model": "{{TEXT_MODEL}}",
  "choices": [
    {
      "index": 0,
      "text": "Most people lose some hair every day. It's a natural process called shedding. On average, you can lose 50-100 strands of hair a day. However, if you notice an increase in hair loss or if you are experiencing hair loss in large clumps, then it could be a sign of an underlying condition. Hair loss can be triggered by a number of factors such as stress, illness, hormone imbalance, medication, and genetics.\nIf you're concerned about hair loss, you should consult with a trichologist, a hair and scalp specialist. They can help determine the cause of your hair loss and recommend the appropriate treatment."
    }
  ]
}
```

### Streaming

**Request**

```json
{
  "model": "{{TEXT_MODEL}}",
  "prompt": "Will I lose my hair?",
  "frequency_penalty": 0.1,
  "logit_bias": {
    "128000": 10
  },
  "max_tokens": 1000,
  "presence_penalty": 0.1,
  "stop": "hello",
  "stream": true,
  "temperature": 1,
  "top_p": 1,
  "top_k": 50,
  "output": {
    "toxicity": true
  },
  "input": {
    "pii": "replace",
    "pii_replace_method": "random"
  }
}
```

**Response**

```json
{
  "object": "text_completion",
  "created": 1750950603,
  "model": "{{TEXT_MODEL}}",
  "choices": [
    {
      "index": 0,
      "text": " please",
      "finish_reason": null,
      "stop_reason": null
    }
  ],
  "id\"": "cmpl-4b78eda0a81145b495fc13c2b3da2041"
}
```