EmbeddedLLM
We are committed to making production-grade AI inference as accessible and reliable as electricity, powered by vLLM.
Pinned Loading
Repositories
Showing 10 of 83 repositories
- agentic-api Public Forked from vllm-project/agentic-api
Stateful API logic for agentic applications using vLLM
- vllm Public Forked from vllm-project/vllm
vLLM: A high-throughput and memory-efficient inference and serving engine for LLMs
- litellm Public Forked from BerriAI/litellm
Python SDK, Proxy Server (LLM Gateway) to call 100+ LLM APIs in OpenAI format - [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, Replicate, Groq]
- TokenVisor-charts Public
- llm-d-router Public Forked from llm-d/llm-d-router
llm-d Router: The intelligent entry point for inference requests
- llm-d Public Forked from llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
- Mooncake Public Forked from kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
- tokenvisor-pi Public
Top languages
Loading…
Most used topics
Loading…