ChatGPT API in Production: A Practical Guide

# ChatGPT API in Production: A Practical Guide

The ChatGPT API isn’t magic—it’s an HTTP endpoint that takes JSON in and gives JSON out. If you can make a network request, you can use it. This guide shows you exactly how to integrate the API into your projects, with real code you can run today.

I’m covering setup, making requests, handling responses, managing costs, and the common mistakes that trip up developers. No fluff, no hype—just what works in 2026.

## Getting Your API Key

Before writing any code, you need an API key. Here’s how:

1. Go to [platform.openai.com](https://platform.openai.com)
2. Sign up or log in
3. Navigate to **API Keys** → **Create new secret key**
4. Copy it immediately—it only shows once

**Never commit your API key to version control.** Add it to your environment:

“`bash
# .env file (add to .gitignore)
OPENAI_API_KEY=sk-your-key-here
“`

You need to add credit to your account. The API works on pay-per-use, not a subscription. Check your usage dashboard regularly—it’s easy to overspend when testing.

## Making Your First Request

The API endpoint is straightforward. Here’s a raw curl request to understand what’s happening under the hood:

“`bash
curl https://api.openai.com/v1/chat/completions \
-H “Content-Type: application/json” \
-H “Authorization: Bearer $OPENAI_API_KEY” \
-d ‘{
“model”: “gpt-4o”,
“messages”: [
{“role”: “user”, “content”: “Say hello in one sentence.”}
],
“max_tokens”: 50
}’
“`

The response looks like this:

“`json
{
“id”: “chatcmpl-xxx”,
“object”: “chat.completion”,
“created”: 1700000000,
“model”: “gpt-4o”,
“choices”: [
{
“index”: 0,
“message”: {
“role”: “assistant”,
“content”: “Hello! How can I help you today?”
},
“finish_reason”: “stop”
}
],
“usage”: {
“prompt_tokens”: 10,
“completion_tokens”: 8,
“total_tokens”: 18
}
}
“`

That’s the core of it—send messages, get a response back.

## Python Integration (The Right Way)

While curl works for testing, your projects need proper libraries. Here’s a real implementation:

“`python
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ[“OPENAI_API_KEY”])

def chat_with_gpt(prompt: str, model: str = “gpt-4o”) -> str:
response = client.chat.completions.create(
model=model,
messages=[
{“role”: “user”, “content”: prompt}
],
max_tokens=500,
temperature=0.7
)

return response.choices[0].message.content

# Usage
result = chat_with_gpt(“Explain async/await in 2 sentences”)
print(result)
“`

That’s the official `openai` Python library. Install it with:

“`bash
pip install openai
“`

The library handles retries, timeouts, and connection management. Don’t write your own HTTP wrapper—use this.

## Streaming Responses for Better UX

Waiting for the full response feels slow. Use streaming to get tokens as they’re generated:

“`python
from openai import OpenAI

client = OpenAI(api_key=os.environ[“OPENAI_API_KEY”])

stream = client.chat.completions.create(
model=”gpt-4o”,
messages=[{“role”: “user”, “content”: “Write a Python function that sorts a list”}],
stream=True
)

for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end=””, flush=True)
“`

This prints tokens incrementally. Essential for chatbots, real-time apps, or anywhere the user expects immediate feedback.

## Cost Management and Rate Limits

The API charges per token. Here’s what matters:

**Token pricing (as of 2026):**
– GPT-4o: ~$2.50/1M input tokens, ~$10/1M output tokens
– GPT-4o-mini: ~$0.15/1M input, ~$0.60/1M output
– o1-preview: Higher pricing, different token handling

**To estimate costs before sending:**

“`python
import tiktoken

def estimate_tokens(text: str, model: str = “gpt-4o”) -> int:
encoding = tiktoken.encoding_for_model(model)
return len(encoding.encode(text))

# Before API call
prompt = “Your long prompt here…”
estimated = estimate_tokens(prompt)
print(f”Estimated cost: ${estimated * 0.0000025:.4f}”)
“`

**Rate limits** depend on your tier. Check the platform dashboard. If you hit limits, implement exponential backoff:

“`python
import time
import openai
from openai import RateLimitError

def make_request_with_retry(prompt: str, max_retries: 3):
for attempt in range(max_retries):
try:
response = client.chat.completions.create(
model=”gpt-4o”,
messages=[{“role”: “user”, “content”: prompt}]
)
return response.choices[0].message.content
except RateLimitError as e:
if attempt == max_retries – 1:
raise e
wait_time = 2 ** attempt
print(f”Rate limited. Waiting {wait_time}s…”)
time.sleep(wait_time)
“`

## Real Example: A Code Review Assistant

Here’s a practical project—a CLI tool that reviews code files:

“`python
#!/usr/bin/env python3
“””Code review assistant CLI tool”””

import os
import sys
from openai import OpenAI

client = OpenAI(api_key=os.environ[“OPENAI_API_KEY”])

SYSTEM_PROMPT = “””You are a senior developer reviewing code.
Provide brief, actionable feedback. Focus on bugs, security issues,
and performance problems. Skip style preferences unless critical.”””

def review_file(filepath: str) -> str:
with open(filepath, ‘r’) as f:
code = f.read()

response = client.chat.completions.create(
model=”gpt-4o”,
messages=[
{“role”: “system”, “content”: SYSTEM_PROMPT},
{“role”: “user”, “content”: f”Review this code:\n\n“`{code}“`”}
],
temperature=0.3
)

return response.choices[0].message.content

if __name__ == “__main__”:
if len(sys.argv) != 2:
print(“Usage: python review.py “)
sys.exit(1)

filepath = sys.argv[1]
print(f”Reviewing {filepath}…\n”)
review = review_file(filepath)
print(review)
“`

Run it:

“`bash
python review.py my_app/utils.py
“`

This is the pattern for most API integrations: read input → send to API → process output. Customize the prompt, change the model, add retries—and you have a production-ready tool.

## Common Pitfalls

**1. Context window limits**
Each model has a maximum context length (e.g., 128K tokens for gpt-4o). If your prompt + history exceeds this, you get an error or truncated responses. Monitor your token count.

**2. JSON mode isn’t magic**
Use `response_format={“type”: “json_object”}` when you need JSON—but the model still generates text that looks like JSON. Validate the output.

**3. Temperature matters**
– `temperature=0`: Deterministic, best for factual tasks
– `temperature=0.7-0.9`: Creative, best for brainstorming
– Higher temperature = more hallucinations

**4. System prompts get ignored**
The system prompt sets behavior, but users often override it with clever jailbreaking. Don’t use the API for security-critical decisions without human oversight.

**5. No built-in memory**
Each request is independent. For conversation history, you must pass the full message array:

“`python
messages = [
{“role”: “system”, “content”: “You are a helpful assistant.”},
{“role”: “user”, “content”: “My name is Bob.”},
{“role”: “assistant”, “content”: “Hello Bob!”},
{“role”: “user”, “content”: “What’s my name?”} # Model knows from context
]
“`

This adds up in costs fast—implement truncation or summarization for long conversations.

## Key Takeaways

– Get your API key from platform.openai.com, never commit it to git
– Use the official `openai` Python library—don’t write raw HTTP wrappers
– Stream responses for better UX in interactive applications
– Monitor token usage and implement cost estimation before API calls
– Implement retry logic with exponential backoff for rate limits

## Next Steps

1. **Create your API key** at platform.openai.com if you haven’t already
2. **Run the code examples** in this article to understand the request/response flow
3. **Build something small**—a script that summarizes a file, answers questions about your codebase, or formats log output
4. **Add cost tracking** by logging token usage from API responses
5. **Read the official docs** at platform.openai.com/docs for streaming, function calling, and advanced features

The API is straightforward once you see it as what it is: an HTTP endpoint that takes text in and gives text out. Build from there.