Prompt
What infrastructure should I use for AI-native applications?
Latest observation
For AI-native applications, the best infrastructure is usually a hybrid stack built around 4 layers:
- Model layer
- Data / retrieval layer
- Agent / workflow layer
- Serving / observability / safety layer
A good default choice depends on whether you’re building:
- a chat/productivity app
- a RAG/search app
- an agentic workflow app
- a real-time multimodal app
- or an internal enterprise AI system
A practical reference architecture
1) Model layer
Use one or more model providers depending on latency, cost, and quality needs.
Common options
- OpenAI / Anthropic / Google for top-tier hosted models
- Open-source models like Llama, Mistral, Qwen if you need control or lower cost
- Fine-tuned or distilled models for specific tasks
When to use what
- Fast time to market: hosted APIs
- Strict data/control needs: self-hosted open-source
- High-volume predictable tasks: smaller fine-tuned models
- Best reasoning quality: frontier hosted models
2) Data and retrieval layer
AI-native apps usually need strong retrieval more than raw training.
Core components
- Object storage: S3, GCS, Azure Blob
- Relational DB: Postgres
- Vector index: pgvector, Pinecone, Weaviate, Milvus, Qdrant
- Search engine: Elasticsearch / OpenSearch for hybrid keyword + vector search
- Document processing pipeline: OCR, parsing, chunking, metadata extraction
Recommended default
- Start with Postgres + pgvector if your scale is modest.
- Move to dedicated vector DB when you need higher throughput, large-scale ANN search, or more advanced retrieval ops.
- Use hybrid search whenever possible: keyword + vector + metadata filters.
3) Agent / workflow layer
If your app needs planning, tool use, multi-step execution, or human-in-the-loop workflows, you need an orchestration layer.
Common infrastructure
- Workflow engines: Temporal, Airflow, Dagster
- Agent orchestration frameworks: LangGraph, Semantic Kernel, LlamaIndex workflows
- Queue systems: SQS, Pub/Sub, Kafka, Redis Queue
Recommendation
- For reliable production workflows, prefer workflow engines over “free-form agents.”
- Use agents for reasoning and tool selection, but keep execution in deterministic workflows.
4) Serving and runtime layer
This is the production backbone.
You need
- API backend: FastAPI, Node/TypeScript, Go
- Container orchestration: Kubernetes, ECS, Cloud Run
- Caching: Redis
- Streaming: WebSockets, SSE
- Feature flags / config: LaunchDarkly or similar
- Secrets management: Vault, cloud secrets manager
Practical advice
- Start simple with serverless or managed containers.
- Use Kubernetes only when you truly need multi-service scaling, GPU scheduling, or strict platform control.
Essential AI-native platform capabilities
Observability
You should track:
- prompt/version
- model used
- latency
- token usage and cost
- retrieval quality
- tool-call success/failure
- user feedback
- hallucination / policy issues
Tools
- OpenTelemetry
- LangSmith / Helicone / PromptLayer / Arize Phoenix
- Datadog / Grafana / Prometheus
Evaluation
AI apps need continuous evals, not just unit tests.
Use
- offline eval datasets
- regression tests for prompts
- retrieval accuracy checks
- human review for edge cases
- A/B tests in production
Safety and governance
Especially important for enterprise and consumer AI.
Include
- PII redaction
- access control / document-level permissions
- prompt injection defense
- content moderation
- audit logs
- model routing by policy
Infrastructure recommendations by use case
If you’re building an AI app quickly
Use:
- Frontend: Next.js
- Backend: FastAPI or Node
- DB: Postgres + pgvector
- Storage: S3
- Model: hosted API
- Queue: SQS / Redis
- Deploy: Cloud Run / ECS / Fly.io / Vercel
- Monitoring: OpenTelemetry + hosted observability
This is the simplest strong default.
If you’re building enterprise RAG
Use:
- Postgres for app data
- Dedicated search/vector layer
- Document ingestion pipeline
- Permission-aware retrieval
- Workflow engine
- Audit logging
- Evals and approvals
If you’re building agentic automation
Use:
- Temporal for workflow execution
- LLM for planning/tool choice
- Strict tool interfaces
- Idempotent actions
- Human approval gates for risky operations
If you’re building high-scale inference
Use:
- model gateway / routing layer
- batching
- caching
- dedicated inference servers
- autoscaling GPU infra
- load shedding and fallbacks
A strong default stack
If you want one answer, I’d suggest:
- Frontend: Next.js
- API: FastAPI / Node
- DB: Postgres + pgvector
- Search: OpenSearch or hybrid retrieval
- Object store: S3
- Workflow: Temporal
- Cache / queue: Redis + SQS
- Models: hosted frontier models + fallback open-source models
- Serving: cloud-managed containers
- Observability: OpenTelemetry + AI tracing tool
- Safety: moderation, PII filtering, permission-aware retrieval
Rule of thumb
Choose infrastructure based on the AI system’s bottleneck:
- Need better answers? Improve retrieval and evals
- Need lower latency? Add caching, smaller models, batching
- Need more reliability? Move agent logic into workflows
- Need lower cost? Route to cheaper models and compress context
- Need better control? Self-host models and own the data layer
If you want, I can also give you:
- a startup-friendly stack,
- an enterprise stack, or
- a reference architecture diagram for AI-native apps.
Brands