Prompt
How do I add chat and text generation to my app with an LLM API?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To add chat and text generation to your app with an LLM API, the basic pattern is:
-
Pick an API provider
- Examples: OpenAI, Anthropic, Google, Azure OpenAI, etc.
- Choose based on model quality, price, latency, tool support, and safety features.
-
Create a backend endpoint in your app
- Don’t call the LLM API directly from the browser/mobile app if it requires a secret API key.
- Your app’s client sends user input to your server, and your server calls the LLM API.
-
Send a prompt or message history
- For text generation, send a prompt like “Write a product description for…”
- For chat, send a list of messages:
- system: instructions/behavior
- user: user’s message
- assistant: previous responses
-
Return the model’s output to the UI
- Display it in your chat window or editor.
- Optionally use streaming so text appears token-by-token for a better UX.
-
Add guardrails
- Validate input length.
- Handle rate limits/errors.
- Filter sensitive content if needed.
- Log usage and monitor costs.
Typical architecture
Client app
- Chat box / prompt input
- Displays generated text
Your backend
- Receives user message
- Adds system instructions
- Calls LLM API
- Streams or returns result
LLM provider
- Generates the response
Minimal example flow
Chat request
Your backend sends something like:
{
"model": "some-model",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Help me write a short email to a client."}
]
}
Text generation request
{
"model": "some-model",
"input": "Write a short email to a client apologizing for the delay."
}
Example implementation idea
Backend route (pseudo-code)
from flask import Flask, request, jsonify
import os
app = Flask(__name__)
@app.post("/api/chat")
def chat():
user_message = request.json["message"]
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": user_message}
]
# Call your LLM provider here
result = llm_client.generate(messages=messages)
return jsonify({"reply": result})
Frontend
async function sendMessage(message) {
const res = await fetch("/api/chat", {
method: "POST",
headers: {"Content-Type": "application/json"},
body: JSON.stringify({ message })
});
const data = await res.json();
displayReply(data.reply);
}
Best practices
- Use a system prompt to define tone and rules.
- Keep conversation history if you want the model to remember context.
- Stream responses for chat-like experiences.
- Trim old messages to stay within token limits.
- Cache or summarize long conversations for cost control.
- Never expose your API key in the frontend.
Features you may want next
- Function calling / tool use: let the model trigger actions like search, database lookup, or sending emails.
- RAG (Retrieval-Augmented Generation): connect the model to your documents.
- Moderation/safety filters
- Multi-turn memory
- Prompt templates for reusable tasks
If you want, I can show you a working example in Node.js, Python, or React using a specific LLM API.