Popular repositories Loading
-
qwen3.5-fp8-cuda
qwen3.5-fp8-cuda PublicQwen3.5-4B FP8 推理引擎 · 手写 CUDA 融合算子 · 54 tok/s on RTX 5060
Python
-
ai-inference-lab
ai-inference-lab PublicLLM inference engineering on a single 8 GB GPU: FP8/INT4-AWQ quantization, a 200-question fairness-controlled quality eval, vLLM tuning, a FastAPI gateway, and a verified 7.49 GiB offline bundle.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.