Prompt
I'm building an internal tool and want the quickest way to add LLM prompts
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
The quickest path is usually:
-
Pick a simple API-first LLM provider
- OpenAI, Anthropic, or a hosted open-source provider.
- Use the provider’s chat/completions API directly.
-
Wrap prompts in a tiny internal service
- Create one endpoint like
POST /llm/run. - Inputs:
prompt_name,variables,context,user_id. - Outputs: raw model response plus metadata.
- Create one endpoint like
-
Store prompts as versioned templates
- Keep them in:
- a YAML/JSON file in your repo, or
- a small DB table if non-devs will edit them.
- Example:
summarize_ticket_v1: system: "You summarize support tickets." user: "Summarize this ticket in 3 bullets: {{ticket_text}}"
- Keep them in:
-
Add a minimal prompt renderer
- Substitute variables like
{{ticket_text}}. - Validate required variables before sending to the model.
- Substitute variables like
-
Log everything
- Prompt version
- Input variables
- Model name
- Output
- Token usage
- Latency This saves a lot of time later.
-
Use a lightweight SDK/framework only if needed
- If you want speed with some structure:
- LangChain: lots of integrations, but can be heavier than needed.
- LlamaIndex: good if retrieval/RAG is a big part.
- Promptfoo: great for testing prompt variants.
- If the goal is “fastest to ship,” many teams just build a thin wrapper themselves.
- If you want speed with some structure:
Recommended MVP architecture
- Frontend/internal tool: lets users pick a prompt and fill variables
- Backend:
- prompt registry
- renderer
- LLM API client
- logger
- Optional: approval/versioning workflow for prompt changes
If you want the absolute fastest start
Use:
- a single provider SDK,
- prompts in YAML,
- one backend endpoint,
- basic logging.
If you want, I can give you:
- a minimal Node.js example,
- a minimal Python/FastAPI example, or
- a prompt registry design for an internal tool.