LLM API Guide
Build production-ready AI applications: API fundamentals, prompt engineering, rate limiting, cost optimization, and scalable patterns for LLM-powered services.
Lessons
Master the practical skills to build and deploy LLM-powered applications at scale.
LLM API Fundamentals +
| Provider | Model | Pricing | Key Features |
|---|---|---|---|
| OpenAI | GPT-4, GPT-3.5 | Pay-per-token | Most mature, extensive docs, function calling |
| Anthropic | Claude | Pay-per-token | Safer, longer context windows |
| Gemini | Pay-per-token | Multimodal, function calling | |
| Cohere | Command | Pay-per-token | Embeddings, reranking |
| DeepSeek | DeepSeek-Chat | Low-cost | Open-source friendly |
Basic API Call Structure
Most LLM APIs follow a similar pattern. Here's a complete example with OpenAI:
import openai, os
client = openai.OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms."}
],
temperature=0.7,
max_tokens=500
)
print(response.choices[0].message.content)Temperature and Sampling
| Parameter | Range | Effect |
|---|---|---|
temperature | 0.0 - 2.0 | Creativity vs. determinism (lower = more predictable) |
top_p | 0.0 - 1.0 | Nucleus sampling for diversity |
max_tokens | 1 - 4096 | Response length limit |
frequency_penalty | 0.0 - 2.0 | Reduce repetitive responses |
Error Handling
from openai import OpenAI, OpenAIError
import time
def call_llm_safely(client, messages, max_retries=3):
for attempt in range(max_retries):
try:
response = client.chat.completions.create(
model="gpt-4o", messages=messages, temperature=0.7
)
return response.choices[0].message.content
except OpenAIError as e:
if attempt < max_retries - 1:
time.sleep(2 ** attempt)
else:
raise eSystem Prompts: Setting the Stage for AI Behavior +
System prompts are the invisible director setting expectations. They run before all conversation, shaping how the model responds. Without them, models default to generic helpfulness. With them, you get consistent, reliable outputs.
Why they matter: imagine asking a junior developer to summarize quarterly earnings. Without context, they'd ask for spreadsheet access. With "summarize email reports from finance@company.com sent before March 31," they ship the right work in minutes. System prompts replace dozens of follow-up questions.
Anatomy of a Good System Prompt
You are a senior ML engineer specializing in inference optimization. Provide concise, production-focused answers. Always explain tradeoffs. Never mention training procedures unless asked. Format code in triple backticks. Respond to technical questions in under 200 words.
This prompt establishes capabilities, constraints, and tone. The model knows its role, limits, and audience.
System Prompts vs User Prompts
A user prompt with "summarize this" gets a summary. A user prompt "summarize using these four headings" still works, but a system prompt "you are a technical writer" makes summaries consistently structured. System prompts set the stage, user prompts provide the script.
Common Patterns
- Role-based: "You are a Python expert..." → direct, practical answers with code examples.
- Constraint-based: "Respond in under 100 words..." → concise output without follow-up length concerns.
- Format-based: "Always use markdown..." → consistent formatting for automated parsing.
- Personality: "Be direct but friendly..." → tone without sacrificing clarity.
Real-World Example: API Client Helper
You are an API integration specialist. Write production-ready Python using httpx and type hints. Include proper error handling and docstrings. Never include API keys. Use placeholders.
User prompt: "Create a client for OpenWeatherMap" → result: code with retry logic, proper type definitions, and clear error messages, without you needing to specify those requirements.
Dynamic System Prompts
Current task: Debug slow database queries Context: PostgreSQL, 1M+ rows, queries over 5 seconds Focus: Query optimization, indexing, EXPLAIN ANALYZE
Now the model shifts from general advice to specific debugging strategies.
Testing System Prompts
Bad system prompts produce inconsistent results. Test by asking the same question 3 times. If answers vary wildly, the prompt lacks clarity or constraints. Good system prompts yield predictable responses across inputs, consistency enables automation.
Practical Exercise
Create a system prompt for code review:
You are a senior Python developer reviewing pull requests. Focus on security, performance, and maintainability. Point out specific lines with line numbers. Suggest concrete alternatives. Remain constructive and professional.
Test with this diff:
- if user_input: + if len(user_input) > 0:
A good system prompt should flag the unnecessary len() call as a performance anti-pattern.
Key takeaways
System prompts are your first line of quality control. They define who the model is, what constraints apply, and how responses should look. Spend time crafting them, they're worth more than perfect user prompts.
Structured Outputs: Turning Freeform Text Into Predictable Data +
LLMs excel at unstructured text but stumble when you need reliable data. Ask for "a list of features" and get prose. Ask for JSON and sometimes still get "Sure, here's the JSON: {...}" wrapped in commentary. Structured outputs turn probabilistic text back into deterministic code.
The Problem: Unreliable Output
User prompt: "Give me three Python features" produces a paragraph of prose. Three features? Hard to parse. Should you extract with regex? What happens when the format changes next time?
The Solution: Constrained Output
[
{"name": "list comprehensions", "description": "..."},
{"name": "decorators", "description": "..."},
{"name": "type hints", "description": "..."}
]The LLM produces valid JSON. Your code parses it. No guesswork.
How Structured Outputs Work
response_format = {"type": "json_object"}
# or, for more control:
response_format = {
"type": "json_schema",
"json_schema": {
"name": "feature_list",
"schema": {
"type": "object",
"properties": {
"features": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"description": {"type": "string"}
},
"required": ["name", "description"]
}
}
}
}
}
}Real-World Example: API Specification
Without structured output, "design a REST API for a todo app" gets you prose you copy-paste and debug. With it:
schema = {
"type": "object",
"properties": {
"endpoints": {
"type": "array",
"items": {
"type": "object",
"properties": {
"method": {"type": "string"},
"path": {"type": "string"},
"description": {"type": "string"},
"body": {"type": "object"}
}
}
}
}
}Output imports directly into your FastAPI routes. Zero manual parsing.
Common Patterns
Batch extraction: pull structured data from unstructured text (e.g. support emails → {"name", "email", "issue"}).
Template validation: ensure the model fills templates correctly.
Configuration generation: turn natural language into config objects, "make this AWS Lambda config" → structured JSON parameters.
Trade-offs
Pros: deterministic output, no parsing logic, type-safe downstream processing.
Cons: rigid format, less natural language flexibility, schema definition effort.
Use structured outputs when you need reliability. Use freeform when you need explanation.
Getting Started
import openai, json
response = openai.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "List 3 Python features as JSON array: {name, description}"}],
response_format={"type": "json_object"}
)
features = json.loads(response.choices[0].message.content)Key takeaways
Structured outputs are guardrails, not cages. They force LLMs into the format you need, eliminating brittle parsing logic. Define your schema once, get perfect data every time.
Function Calling: Making LLMs Call Your Code +
LLMs are great at conversation but terrible at running code. Function calling bridges this gap, it lets the model decide what to do, while you write how to do it.
Without function calling: you write "call search function with this query" in your prompt, the model hallucinates search parameters, your code doesn't run. With function calling: the model returns structured function arguments and your code executes them directly.
How It Works
You provide two things to the API: function definitions (the tools available) and messages (the conversation). The model replies with either text or a function call object. When it calls a function, you execute it and send the result back, the model continues reasoning.
Simple Example: Weather Lookup
functions = [{
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. San Francisco"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "default": "celsius"}
},
"required": ["city"]
}
}]
# User asks: "What's the weather in Tokyo?"
# Model responds:
# {"name": "get_weather", "arguments": "{\"city\": \"Tokyo\", \"unit\": \"celsius\"}"}
def get_weather(city: str, unit: str = "celsius") -> str:
return f"{city}: 22°{unit[0].upper()}"Complex Example: Restaurant Booking
functions = [{
"name": "book_table",
"description": "Book a restaurant reservation",
"parameters": {
"type": "object",
"properties": {
"restaurant": {"type": "string"},
"party_size": {"type": "integer", "minimum": 1, "maximum": 20},
"datetime": {"type": "string", "format": "date-time"},
"occasion": {"type": "string", "enum": ["birthday", "anniversary", "date", "none"]}
},
"required": ["restaurant", "party_size", "datetime"]
}
}]
# "Book a table for 4 at Sushi House tomorrow at 7 PM for my anniversary" ->
# {"name": "book_table", "arguments": {"restaurant": "Sushi House", "party_size": 4,
# "datetime": "2024-01-15T19:00:00", "occasion": "anniversary"}}When to Use Function Calling
Use it when: the model needs to make decisions, external data is required, you need reliable input extraction, or it's part of an automation workflow.
Don't use it when: it's pure conversation, a simple calculation the model can just reason through, or the information is already in context.
Building a Complete Loop
import openai, json
available_functions = {
"get_weather": lambda city, unit="celsius": f"{city}: 22°",
"book_table": lambda restaurant, party_size, datetime, occasion="none": "Booked!"
}
def run_conversation(messages):
while True:
response = openai.chat.completions.create(
model="gpt-4o", messages=messages, functions=[...], function_call="auto"
)
message = response.choices[0].message
if message.get("function_call"):
func_name = message["function_call"]["name"]
func_args = json.loads(message["function_call"]["arguments"])
result = available_functions[func_name](**func_args)
messages.append({"role": "assistant", "content": result, "function_call": message["function_call"]})
else:
return message["content"]
messages = [{"role": "user", "content": "What's the weather in London?"}]
print(run_conversation(messages))Common Patterns
- Orchestrator pattern: LLM decides which functions to call and in what order, great for multi-step workflows.
- Extract pattern: LLM extracts structured data from user input into function arguments, clean separation of NLU and execution.
- Validation pattern: model calls validation functions to check if proposed solutions work.
Tool Use vs Function Calling
Function calling is simpler but less flexible. Tool use (used by agents) supports multiple parallel calls, tool descriptions, and more complex interactions. Start with function calling, move to tool use when you need multi-step reasoning.
Getting Started Checklist
- Define function signatures for your APIs
- Map model calls to actual implementations
- Handle errors (invalid arguments, missing functions)
- Return results in conversational format
- Test with ambiguous and complex queries
Key takeaways
Function calling turns LLMs into decision engines while keeping your code execution reliable. Define clear interfaces, handle the loop, and let the model orchestrate your tools.