Prompt
How do I use an LLM API for chat in a web app?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To use an LLM API for chat in a web app, the usual setup is:
- Frontend collects user messages.
- Backend sends those messages to the LLM API.
- Backend returns the model’s reply to the frontend.
- Optionally, stream tokens for a typing effect.
Basic flow
1) Frontend
You keep a chat history in state and send the latest message to your server.
2) Backend
Your server:
- authenticates with the LLM provider using an API key
- sends the conversation history
- receives the assistant response
- returns it to the browser
Important: don’t call the LLM API directly from the browser if it requires a secret key, because users could steal it.
Example architecture
Browser UI -> Your backend /api/chat -> LLM provider API
Browser UI <- Your backend response <- LLM reply
Example: frontend (simple fetch)
<input id="msg" placeholder="Type a message" />
<button id="send">Send</button>
<div id="chat"></div>
<script>
const chat = document.getElementById("chat");
const msg = document.getElementById("msg");
document.getElementById("send").onclick = async () => {
const userMessage = msg.value;
chat.innerHTML += `<div><b>You:</b> ${userMessage}</div>`;
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
messages: [
{ role: "user", content: userMessage }
]
})
});
const data = await res.json();
chat.innerHTML += `<div><b>Bot:</b> ${data.reply}</div>`;
msg.value = "";
};
</script>
Example: backend (Node.js / Express)
import express from "express";
const app = express();
app.use(express.json());
app.post("/api/chat", async (req, res) => {
const { messages } = req.body;
const response = await fetch("https://api.example-llm.com/v1/chat", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${process.env.LLM_API_KEY}`,
},
body: JSON.stringify({
model: "your-model-name",
messages,
}),
});
const data = await response.json();
res.json({ reply: data.output });
});
app.listen(3000, () => console.log("Server running on 3000"));
Chat history
For a real chat app, send the conversation so the model has context:
[
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Hi" },
{ role: "assistant", content: "Hello! How can I help?" },
{ role: "user", content: "Explain CSS grid." }
]
Most LLM APIs expect messages in this style.
Streaming responses
If you want the reply to appear gradually:
- use Server-Sent Events (SSE), WebSockets, or fetch streaming
- have the backend forward chunks from the LLM API to the client
This improves UX for long responses.
Good practices
- Keep API keys on the server
- Add rate limiting to prevent abuse
- Validate input
- Store conversation state if needed
- Use moderation/filtering for safety if your app is public
- Handle errors gracefully and show retry UI
Minimal implementation checklist
- Build a chat UI
- Create
/api/chaton your server - Store LLM API key in environment variables
- Send conversation messages to the LLM API
- Return assistant output to frontend
- Optionally add streaming
If you want, I can give you:
- a full working example in React + Node.js, or
- an example using a specific provider like OpenAI, Anthropic, or Google Gemini.