Cost-Aware Model Routing Gateway for LLM Applications
Orange
A LLMRouter & OpenAI-compatible inference gateway that intelligently routes requests to the cheapest capable AI model using local embedding-based complexity scoring, semantic caching, and budget-aware failover—optimizing cost without compromising response quality.
- Python
- FastAPI
- LangChain
- FAISS
- Ollama
- OpenAI API
- Embeddings
- Docker




