LLMKube

4周前发布 36 0 0

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Meta...

收录时间:
2026-08-14
LLMKubeLLMKube

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API. (Source Code) Apache-2.0 Go/Docker/K8S

数据统计

相关导航