Prompt
How do I add chat to my app with an LLM API?
Latest observation
To add chat to your app with an LLM API, you generally need 4 pieces:
- A UI for messages
- A backend endpoint that talks to the LLM
- Conversation state/history
- Streaming or polling so replies feel fast
Here’s a practical blueprint.
1) Build the chat UI
Your frontend should let the user:
- type a message
- see their message appear immediately
- show the assistant’s response
- optionally show typing/streaming state
Typical structure:
- message list
- text input
- send button
- loading indicator
2) Send messages to your backend, not directly to the LLM
Don’t call the LLM API from the browser/mobile client unless the API key is safe there.
Instead:
Frontend → Your server → LLM API
Your server:
- validates the request
- attaches system instructions
- adds conversation history
- calls the LLM
- returns the result
3) Keep conversation history
Most chat LLMs need the prior messages so they can respond in context.
A typical message format is:
[
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Hi" },
{ "role": "assistant", "content": "Hello! How can I help?" },
{ "role": "user", "content": "What can you do?" }
]
You usually store this:
- in memory for a quick prototype
- in a database for real apps
Common strategy:
- keep the full conversation in your DB
- send only the most recent N messages to the model to control cost/latency
4) Call the LLM API
Most LLM APIs accept something like:
- model name
- messages
- optional temperature / max tokens
- streaming flag
Example backend flow
- User sends message
- Backend loads conversation history
- Backend appends new user message
- Backend calls LLM API
- Backend stores assistant reply
- Backend sends reply to frontend
5) Use streaming for a better UX
Instead of waiting for the full answer, stream tokens as they generate.
This makes the app feel much faster.
Frontend:
- opens an SSE or WebSocket connection, or uses fetch streaming
- appends text as chunks arrive
Backend:
- forwards streamed output from the LLM API to the client
If you’re building a simple MVP, you can start with non-streaming and add streaming later.
6) Add a system prompt
A system prompt defines behavior and style.
Example:
You are a concise, friendly support assistant. Ask clarifying questions when needed. If you don't know something, say so.
Use it to control:
- tone
- persona
- boundaries
- domain expertise
7) Handle edge cases
You’ll want to handle:
- empty input
- API timeouts
- rate limits
- model errors
- very long conversations
- moderation / safety filtering if needed
If the conversation gets too long, you may need:
- truncation
- summarization
- retrieval from memory/DB
- tool/function calling for external actions
8) Minimal example architecture
Frontend
- renders chat
- POSTs
{ conversationId, message }to your server
Backend
POST /chat- loads conversation history
- calls LLM API
- returns assistant reply
Database
conversationsmessages
9) Pseudocode example
Backend example
app.post("/chat", async (req, res) => {
const { conversationId, message } = req.body;
const history = await loadMessages(conversationId);
const messages = [
{ role: "system", content: "You are a helpful assistant." },
...history,
{ role: "user", content: message }
];
const response = await llmClient.chat.completions.create({
model: "your-model",
messages
});
const reply = response.choices[0].message.content;
await saveMessage(conversationId, "user", message);
await saveMessage(conversationId, "assistant", reply);
res.json({ reply });
});
10) Production considerations
For a real app, also think about:
- Auth: who is allowed to chat
- Rate limiting: prevent abuse
- Logging: monitor errors and usage
- Cost control: cap tokens and history
- Privacy: avoid storing sensitive data unnecessarily
- Safety: moderation and policy checks
- UX: retry button, stop generation, typing indicator
11) Recommended implementation path
If you want the fastest path:
- Build a basic chat UI
- Add a backend
/chatendpoint - Call the LLM with system + user message
- Store conversation history
- Add streaming
- Add auth, rate limits, and persistence
If you want, I can give you:
- a Node.js/Express example
- a Python/FastAPI example
- a React chat UI example
- or a full end-to-end app template