Prompt
How do I build a custom assistant feature with a hosted model API?
Latest observation
To build a custom assistant feature with a hosted model API, think in terms of four layers:
- User interface — chat box, sidebar, workflow button, etc.
- Orchestration layer — your backend that calls the model API, stores context, applies rules, and handles tools.
- Model layer — the hosted LLM API.
- Data/tools layer — your docs, database, search, CRM, calendar, etc.
1) Define the assistant’s job
Be specific about what the assistant should do.
Examples:
- Answer support questions from your docs
- Draft emails based on CRM data
- Summarize meeting notes
- Help users navigate your product
Write down:
- What inputs it gets
- What outputs it should produce
- What it should never do
- Whether it needs tools or just text generation
2) Pick an API pattern
Most hosted model APIs support one or more of these:
A. Simple chat completion
Use this if you just need a conversational assistant.
- Send system instructions
- Send recent conversation history
- Receive a generated response
B. Retrieval-augmented generation (RAG)
Use this if the assistant needs your company knowledge.
- Index docs in a vector database
- Retrieve relevant passages at runtime
- Inject them into the prompt
C. Tool/function calling
Use this if the assistant needs to take action.
- Model decides when to call tools
- Your backend executes the tool
- Return tool results back to the model
- Model writes the final answer
3) Build the assistant loop
Typical flow:
- User sends a message
- Your backend loads conversation state
- Your backend adds:
- system prompt
- relevant memory
- retrieved knowledge
- available tool definitions
- Call the hosted model API
- If the model requests a tool, run it
- Send tool results back to the model
- Return final response to the user
4) Design the system prompt
Your system prompt should define:
- Role and tone
- Scope
- Safety limits
- Formatting rules
- Tool usage rules
Example:
You are a customer support assistant for Acme SaaS.
Answer using only the provided knowledge base and tools.
If you do not know, say you are not sure and offer to escalate.
Be concise and friendly.
5) Add memory carefully
There are usually two kinds of memory:
Short-term memory
Recent conversation turns sent with each request.
Long-term memory
Persistent user preferences or facts stored in your DB.
Don’t send the entire chat history forever. Summarize older context when needed.
6) Add retrieval for private knowledge
If your assistant needs internal docs:
- Split documents into chunks
- Create embeddings
- Store them in a vector DB
- At query time, fetch top-k relevant chunks
- Include those chunks in the model input
This helps with:
- Product docs
- Policies
- Knowledge bases
- Legal/HR documents
7) Add tools for actions
Examples of tools:
- Search customer account
- Create support ticket
- Fetch order status
- Book a meeting
- Send an email
Best practice:
- Keep tools narrow and deterministic
- Validate all tool inputs
- Log every tool call
- Require confirmation before destructive actions
8) Handle reliability
Add guardrails for:
- Rate limits
- Timeouts
- Retries
- Fallback responses
- Output validation
- Prompt injection resistance
For example:
- Never let retrieved text override system instructions
- Sanitize tool inputs
- Use allowlisted tools only
9) Evaluate quality
Test with realistic scenarios:
- Happy path questions
- Ambiguous questions
- Missing data
- Adversarial prompts
- Tool failures
Track:
- Answer accuracy
- Tool success rate
- Latency
- Cost per conversation
- User satisfaction
10) Deploy iteratively
Start with:
- One use case
- One model
- A few tools
- Limited users
Then expand based on logs and feedback.
Minimal architecture example
Frontend chat UI
-> Your backend API
-> Retrieve memory/docs/tools
-> Hosted model API
-> Optional tool execution
-> Hosted model API
<- Final answer
Practical implementation stack
A common stack looks like:
- Frontend: React, Next.js, mobile app
- Backend: Node.js, Python, Go
- Model API: hosted LLM provider
- Retrieval: Pinecone, Weaviate, pgvector, Elasticsearch
- Storage: Postgres, Redis
- Observability: logs, tracing, prompt/version tracking
If you want a fast MVP
Start with:
- A system prompt
- Recent chat history
- One retrieval source
- One or two tools
- Basic logging
That gets you a useful assistant without overengineering.
If you want, I can also show:
- a sample backend flow in Python or Node.js,
- a prompt template, or
- a tool-calling example end to end.