Skip to main content

Chat Completions

Basic Request

With Parameters

Streaming Responses

Stream tokens as they’re generated:

Streaming with Event Handling

Function Calling

Use function calling for structured outputs:

Multi-turn Conversations

Maintain conversation history:

Response Object

Structure

Finish Reasons

Advanced Parameters

Temperature and Sampling

Top-p Sampling

Presence and Frequency Penalties

Vision Models

Send images for analysis:

Batch Inference

Process multiple prompts efficiently:

Stop Sequences

Define custom stop sequences:

Logit Bias

Bias token selection:

Response Caching

Reduce costs with response caching:

Error Handling

Next Steps

Code Examples

Complete code examples and patterns

Pipelines

Orchestrate complex workflows