← OSS.Radar home

Powered by aegismemory.com · Aegis Memory repository

Model Serving AI repositories

OSS Radar projects in the model serving category.

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Category
model serving
Stars
88,472
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +669 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: 30-day commits, lifetime contributors, response activity.

amd blackwell cuda deepseek deepseek-v3 gpt

ray-project/ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

Category
model serving
Stars
43,468
Readiness
ready (97/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +68 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals; open issue backlog is stable or shrinking

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: 30-day commits, lifetime contributors, response activity.

data-science deep-learning deployment distributed hyperparameter-optimization hyperparameter-search

gitleaks/gitleaks

Find secrets with Gitleaks 🔑

Category
model serving
Stars
28,527
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +116 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-powered ci-cd cicd cli data-loss-prevention devsecops

meta-llama/llama-cookbook

Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model f

Category
model serving
Stars
18,554
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +10 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai finetuning langchain llama llama2 llm

tensorzero/tensorzero

TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

Category
model serving
Stars
11,726
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
70/100 · low confidence

Why: +38 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-engineering anthropic artificial-intelligence deep-learning genai

mistralai/mistral-inference

Official inference library for Mistral models

Category
model serving
Stars
10,838
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
70/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: repository is archived. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference mistralai

Tiiny-AI/PowerInfer

High-speed Large Language Model Serving for Local Deployment

Category
model serving
Stars
9,704
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +15 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

large-language-models llama llm llm-inference local-inference

FedML-AI/FedML

FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on a

Category
model serving
Stars
4,055
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 283 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agent deep-learning distributed-training edge-ai federated-learning inference-engine

Xiangyue-Zhang/auto-deep-researcher-24x7

🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Worker architecture, constant-size memory.

Category
model serving
Stars
1,260
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +33 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agent autonomous-agent claude-code deep-learning experiment-automation gpu

FellouAI/eko

Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai

Category
model serving
Stars
4,948
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
26/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 157 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agentic-ai agentic-ai-development agentic-framework agentic-workflow agents

superduper-io/superduper

Superduper: End-to-end framework for building custom AI applications and agents.

Category
model serving
Stars
5,310
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 340 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai chatbot data database distributed-ml inference

decodingai-magazine/llm-twin-course

🤖 𝗟𝗲𝗮𝗿𝗻 for 𝗳𝗿𝗲𝗲 how to 𝗯𝘂𝗶𝗹𝗱 an end-to-end 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗿𝗲𝗮𝗱𝘆 𝗟𝗟𝗠 & 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 using 𝗟𝗟𝗠𝗢𝗽𝘀 best practices: ~ 𝘴𝘰𝘶𝘳𝘤𝘦 𝘤𝘰𝘥𝘦 + 12 𝘩𝘢𝘯𝘥𝘴-𝘰𝘯 𝘭𝘦𝘴𝘴𝘰𝘯𝘴

Category
model serving
Stars
4,382
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aws bytewax comet-ml course docker generative-ai

intel/intel-extension-for-transformers

⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡

Category
model serving
Stars
2,175
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 668 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

4-bits autoround chatbot chatpdf gaudi3 habana

jia-gao/leanctx

Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code. Python SDK, LLMLingua-2, MIT.

Category
model serving
Stars
316
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

anthropic cost-optimization gemini langchain langgraph llm

bentoml/OpenLLM

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

Category
model serving
Stars
12,452
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +10 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bentoml fine-tuning llama llama2 llama3-1 llama3-2

vllm-project/semantic-router

A programmable Mixture-of-Models router for heterogeneous LLM inference

Category
model serving
Stars
5,126
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +35 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-gateway bert-classification fine-tuning golang huggingface-candle huggingface-transformers

predibase/lorax

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

Category
model serving
Stars
3,823
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

fine-tuning gpt llama llm llm-inference llm-serving

openvinotoolkit/openvino

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

Category
model serving
Stars
10,619
Readiness
ready (99/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +27 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai computer-vision deep-learning deploy-ai diffusion-models generative-ai

Netflix/metaflow

Build, Manage and Deploy AI/ML Systems

Category
model serving
Stars
10,208
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai aws azure cost-optimization datascience

bentoml/BentoML

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

Category
model serving
Stars
8,766
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +20 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-inference deep-learning generative-ai inference-platform llm llm-inference

evidentlyai/evidently

Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

Category
model serving
Stars
7,792
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +21 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-drift data-quality data-science data-validation generative-ai hacktoberfest

alvinreal/awesome-opensource-ai

Curated list of the best truly open-source AI projects, models, tools, and infrastructure. Daily updated.

Category
model serving
Stars
4,444
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +53 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai artificial-intelligence awesome awesome-list generative-ai

jina-ai/serve

☁️ Build multimodal AI applications with cloud-native stack

Category
model serving
Stars
21,864
Readiness
needs review (45/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
risky
Maintenance risk
76/100 · high confidence

Why: +2 stars in 7 days; 100+ lifetime contributors

Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking

Strongest signals: contributor breadth, issue load, documentation. Risks: no push in 501 days, latest release is 633 days old, 27 open issues with no maintainer issue responses in 30 days, no pull-request review responses in 30 days, no maintainer response activity in 30 days. Missing inputs: None.

Capped lower bounds: lifetime contributors.

cloud-native cncf deep-learning docker fastapi framework

bricks-cloud/BricksLLM

🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI,

Category
model serving
Stars
1,221
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 579 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai anthropic api artificial-intelligence azure docker

codeproject/CodeProject.AI-Server

CodeProject.AI Server is a self contained service that software developers can include in, and distribute with, their applications in order to augment their apps with the power of AI.

Category
model serving
Stars
973
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 389 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence generative-ai mlops object-detection onnx python

zylon-ai/private-gpt

Complete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference server.

Category
model serving
Stars
57,416
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +26 stars in 7 days; 30 commits in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity; open issue backlog is stable or shrinking

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

ai ai-tools on-premise

suncloudsmoon/awesome-open-source-ai

A curated list of useful open-source AI resources

Category
model serving
Stars
303
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
23/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 557 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-tools embeddings llm-inference llms open-source

liguodongiot/llm-action

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

Category
model serving
Stars
24,869
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +34 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference llm-serving llm-training llmops

Lightning-AI/litgpt

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

Category
model serving
Stars
13,611
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai artificial-intelligence deep-learning large-language-models llm llm-inference

InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Category
model serving
Stars
7,996
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

codellama cuda-kernels deepspeed fastertransformer internlm llama

ai-dynamo/dynamo

A Datacenter Scale Distributed Inference Serving Framework

Category
model serving
Stars
7,700
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +55 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

diffusion disaggregated-serving kubernetes llm-inference omni routing-engine

algorithmicsuperintelligence/openevolve

Open-source implementation of AlphaEvolve

Category
model serving
Stars
6,871
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +48 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

alpha-evolve alphacode alphaevolve coding-agent deepmind deepmind-lab

flashinfer-ai/flashinfer

FlashInfer: Kernel Library for LLM Serving

Category
model serving
Stars
6,128
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +55 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

attention cuda distributed-inference gpu jit large-large-models

kserve/kserve

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

Category
model serving
Stars
5,775
Readiness
ready (99/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +19 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence cncf genai hacktoberfest istio k8s

gpustack/gpustack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

Category
model serving
Stars
5,455
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +33 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ascend cuda deepseek distributed-inference genai high-performance-inference

xlite-dev/Awesome-LLM-Inference

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Category
model serving
Stars
5,446
Readiness
ready (74/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +13 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awesome-llm deepseek deepseek-r1 deepseek-v3 flash-attention flash-attention-3

ruvnet/RuVector

RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.

Category
model serving
Stars
4,408
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-ocr attention-mechanism gnn gnn-model gnns

NVIDIA/GenerativeAIExamples

Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

Category
model serving
Stars
4,143
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gpu-acceleration large-language-models llm llm-inference microservice nemo

FareedKhan-dev/kimi-k3-in-c

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

Category
model serving
Stars
3,299
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

avx2 c99 cpu-inference deep-learning from-scratch inference-engine

spiceai/spiceai

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

Category
model serving
Stars
3,057
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence data data-federation developers full-text-search infrastructure

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

Category
model serving
Stars
1,511
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deepseek glm inference inference-engine large-language-models llm-inference

jax-ml/scaling-book

Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs

Category
model serving
Stars
1,320
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +18 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

jax llm-inference llms roofline tpus

lean-dojo/LeanCopilot

LLMs as Copilots for Theorem Proving in Lean

Category
model serving
Stars
1,308
Readiness
ready (78/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

formal-mathematics lean lean4 llm llm-inference machine-learning

ovg-project/kvcached

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Category
model serving
Stars
1,129
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

elastic-kvcache gpu-mutiplexing gpu-sharing inference-engine kvcache kvcache-optimization

jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

Category
model serving
Stars
1,021
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +53 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai attention batching course cpp cuda

dalisoft/awesome-hosting

List of awesome hosting sorted by minimal plan price

Category
model serving
Stars
921
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai clawdbot cloud database deepseek-r1 free

foldl/chatllm.cpp

Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)

Category
model serving
Stars
915
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference

tingaicompass/AI-Compass

“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。

Category
model serving
Stars
895
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent ai llm llm-inference llm-training nlp

harleyszhang/llm_note

LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

Category
model serving
Stars
887
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda-programming kv-cache llm llm-inference transformer-models triton-kernels

kubernetes-sigs/lws

LeaderWorkerSet: An API for deploying a group of pods as a unit of replication

Category
model serving
Stars
782
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm-inference sig-apps

Avarok-Cybersecurity/atlas

Pure Rust Inference Engine

Category
model serving
Stars
633
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda dgx dgx-spark gb10 llm-inference mamba

pegainfer-project/pegainfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Category
model serving
Stars
631
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda cuda-kernels deepseek gpu inference inference-engine

timmyy123/LLM-Hub

Local AI Assistant

Category
model serving
Stars
535
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +12 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai gemma3 gemma3n gemma4 gemma4-agent-skills gptoss

brontoguana/krasis

Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

Category
model serving
Stars
508
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +13 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpu-inference gguf-model-support gpu-inference high-performance-inference hybrid-inference inference-engine

warpfront/hipfire

RDNA-native LLM inference engine in Rust.

Category
model serving
Stars
496
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +4 stars in 7 days; 40 commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: response activity.

amd-gpu gpu-computing hip llm-inference machine-learning quantization

ome-projects/ome

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

Category
model serving
Stars
487
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +5 stars in 7 days; 39 commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals; open issue backlog is stable or shrinking

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

deepseek k8s kimi-k2 llama llm llm-inference

expectedparrot/edsl

Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs.

Category
model serving
Stars
485
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +1 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: 30-day commits.

anthropic data-labeling deepinfra domain-specific-language experiments llama2

jaylfc/taOS

Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full

Category
model serving
Stars
476
Readiness
needs review (70/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +10 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity

Strongest signals: push recency, commit activity, documentation. Risks: maintenance is concentrated in one contributor. Missing inputs: None.

Capped lower bounds: 30-day commits, response activity.

agent-framework ai-agents ai-platform apple-silicon data-sovereignty distributed-computing

NPC-Worldwide/incognide

Explore the unknown, build the future, own your data.

Category
model serving
Stars
461
Readiness
ready (79/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
8/100 · high confidence

Why: +2 stars in 7 days; 33 commits in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity; open issue backlog is stable or shrinking

Strongest signals: push recency, commit activity, issue load. Risks: no pull-request review responses in 30 days. Missing inputs: None.

agents ai ai-agents artificial-intelligence ide llm-inference

modular/llm-inference-handbook

Everything you need to know about LLM inference

Category
model serving
Stars
376
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +10 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

inference-handbook inference-infrastructure inference-optimization llm llm-inference

jjiantong/Awesome-KV-Cache-Optimization

[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

Category
model serving
Stars
372
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai computer-architecture kv-cache llm llm-inference llm-serving

patchy631/time-to-first-token

A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

Category
model serving
Stars
343
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

learning-resources llm llm-inference mlops roadmap sglang

EfficientMoE/MoE-Infinity

PyTorch library for cost-effective, fast and easy serving of MoE models.

Category
model serving
Stars
338
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

huggingface inference-engine large-language-models llm-inference mixture-of-experts pytorch

Picovoice/picollm

On-device LLM Inference Powered by X-Bit Quantization

Category
model serving
Stars
316
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

compression efficient-inference gemma generative-ai language-model language-models

mlco2/ecologits

🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models through APIs.

Category
model serving
Stars
312
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

genai generative-ai green-ai green-software llm llm-inference

sophgo/LLM-TPU

Run generative AI models in sophgo BM1684X/BM1688

Category
model serving
Stars
296
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bm1684x bm1688 generative-ai internvl3 large-language-models llama3

AstraNetLab/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.

Category
model serving
Stars
292
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

knowledge-injection kvcache kvcache-reuse llm llm-inference llm-task-scheduling

matrixhub-ai/matrixhub

An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.

Category
model serving
Stars
258
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence dynamo huggingface kubernetes llm llm-inference

kalavai-net/kalavai-client

Aggregates compute from spare GPU capacity

Category
model serving
Stars
220
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai free infrastructure llm llm-inference open-source

OpenMachine-ai/transformer-tricks

A collection of tricks and tools to speed up transformer models

Category
model serving
Stars
220
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai arxiv arxiv-papers llm llm-inference llmops

cgbur/llama2.zig

Inference Llama 2 in one file of pure Zig

Category
model serving
Stars
217
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llama llama2 llm llm-inference simd zig

KVignesh122/AssetNewsSentimentAnalyzer

A sentiment analyzer package for financial assets and securities utilizing GPT models.

Category
model serving
Stars
197
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

commodity-trading financial-analysis forex-trading google-search-api investment-analysis llm-inference

feichai0017/loom-infer

Rust-native GPU operator library for LLM inference, built with cuda-oxide

Category
model serving
Stars
186
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda cuda-kernels cuda-oxide gpu high-performance-computing llm-inference

harleyszhang/lite_llama

A light llama-like llm inference framework based on the triton kernel.

Category
model serving
Stars
186
Readiness
ready (79/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

attention llama llama3 llava-llama3 llm llm-inference

Shannon-Data/ShannonBase

The Next-Gen Database for AI—an infrastructure designed for data and AI. As the MySQL of the AI era.

Category
model serving
Stars
175
Readiness
ready (79/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-native ai-native embedding-vectors gpu-acceleration htap in-memory-column-storage

webgptorg/promptbook

Turn your company's scattered knowledge into AI ready Books ✨

Category
model serving
Stars
168
Readiness
needs review (64/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: strong signals despite limited visibility

Strongest signals: push recency. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

autogpt llm-inference openai

codelion/pts

Pivotal Token Search

Category
model serving
Stars
156
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

dataset-generation direct-preference-optimization dpo llm llm-inference llm-steering

Scottcjn/ram-coffers

NUMA-distributed weight banking for LLM inference on IBM POWER8. 147 t/s (8.8x stock). Part of the Proof of Physical AI stack.

Category
model serving
Stars
155
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-inference depin hebbian llama-cpp llm llm-inference

vitalops/datatune

Agentic data transformation on infinite amounts of data

Category
model serving
Stars
149
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

dataops datascience llm llm-inference machinelearning

AmanPriyanshu/Awesome-AI-For-Security

A curated list of tools, papers, and datasets for applying AI to cybersecurity tasks. This list primarily focuses on modern AI technologies like Large Language Models (LLMs), Agents, and Multi-Modal s

Category
model serving
Stars
148
Readiness
ready (75/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agents ai ai-for-security awesome awesome-list

nomic-ai/gpt4all

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

Category
model serving
Stars
77,413
Readiness
high risk (44/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
risky
Maintenance risk
50/100 · high confidence

Why: +13 stars in 7 days; 100+ lifetime contributors

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: contributor breadth, issue load, license. Risks: no push in 437 days, latest release is 529 days old, no pull-request review responses in 30 days. Missing inputs: None.

Capped lower bounds: lifetime contributors.

ai-chat llm-inference

neuralmagic/deepsparse

Sparsity-aware deep learning inference runtime for CPUs

Category
model serving
Stars
3,158
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: repository is archived, no push in 431 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

computer-vision cpus deepsparse inference llm-inference machinelearning

b4rtaz/distributed-llama

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

Category
model serving
Stars
3,031
Readiness
needs review (70/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

distributed-computing distributed-llm llama2 llama3 llm llm-inference

FasterDecoding/Medusa

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

Category
model serving
Stars
2,762
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +9 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 773 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference

SafeAILab/EAGLE

Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).

Category
model serving
Stars
2,500
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
28/100 · low confidence

Why: +17 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load. Risks: no push in 168 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

large-language-models llm-inference speculative-decoding

microsoft/aici

AICI: Prompts as (Wasm) Programs

Category
model serving
Stars
2,074
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 562 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai inference language-model llm llm-framework llm-inference

liltom-eth/llama2-webui

Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.

Category
model serving
Stars
1,938
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 868 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llama-2 llama2 llm llm-inference

aigrantsindia/aigrantsindia

The non profit fostering AI in India through credits, grants, resources

Category
model serving
Stars
1,376
Readiness
high risk (37/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
17/100 · low confidence

Why: +35 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 424 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai grants hackathons india llm-inference llms

character-ai/prompt-poet

Streamlines and simplifies prompt design for both developers and non-technical users with a low code approach.

Category
model serving
Stars
1,154
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
29/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 176 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference prompt prompt-design prompt-engineering prompt-tuning

shyamsaktawat/OpenAlpha_Evolve

OpenAlpha_Evolve is an open-source Python framework inspired by the groundbreaking research on autonomous coding agents like DeepMind's AlphaEvolve.

Category
model serving
Stars
1,046
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 433 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

alphacode alphafold coding-agent discovery distributed-evolutionary-algorithms evolution-computing

zhihu/ZhiLight

A highly optimized LLM inference acceleration engine for Llama and its variants.

Category
model serving
Stars
908
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
24/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 142 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda deepseek-r1 gpt inference-engine llama llm

stanford-mast/blast

Open-source VMs-as-a-service

Category
model serving
Stars
779
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents browser-automation llm-inference python

ghimiresunil/LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing

LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for custom training and inferencing.

Category
model serving
Stars
731
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bert huggingface large-language-models llm-inference llm-training llm-tutorials

feifeibear/long-context-attention

USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

Category
model serving
Stars
685
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

attention-is-all-you-need deepspeed-ulysses llm-inference llm-training pytorch ring-attention

zeux/calm

CUDA/Metal accelerated language model inference

Category
model serving
Stars
646
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: repository is archived, no push in 435 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda llm-inference ml

andrewkchan/yalm

Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O

Category
model serving
Stars
592
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 328 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpp cuda inference-engine llama llamacpp llm

NotPunchnox/rkllama

Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )

Category
model serving
Stars
583
Readiness
needs review (66/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai client client-server ia llm llm-apps

devflowinc/uzi

CLI for running large numbers of coding agents in parallel with git worktrees

Category
model serving
Stars
581
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 429 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai ai codegen go golang llm

rohan-paul/LLM-FineTuning-Large-Language-Models

LLM (Large Language Model) FineTuning

Category
model serving
Stars
577
Readiness
needs review (49/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, fork interest, documentation. Risks: no push in 493 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gpt-3 gpt3-turbo large-language-models llama2 llm llm-finetuning

zjhellofss/KuiperLLama

校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。

Category
model serving
Stars
557
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 283 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpp cuda inference-engine llama2 llama3 llm

yassa9/qwen600

Static suckless single batch CUDA-only qwen3-0.6B mini inference engine

Category
model serving
Stars
555
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 333 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda cuda-programming gpu llamacpp llm llm-inference

microsoft/sarathi-serve

A low-latency & high-throughput serving engine for LLMs

Category
model serving
Stars
515
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 211 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llama llm-inference pytorch transformer

Chen-zexi/vllm-cli

A command-line interface tool for serving LLM using vLLM.

Category
model serving
Stars
507
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 194 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference llm-tools vllm

vectorch-ai/ScaleLLM

A high-performance inference system for large language models, designed for production environments.

Category
model serving
Stars
500
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 231 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda efficiency gpu inference llama llama3

anarchy-ai/LLM-VM

irresponsible innovation. Try now at https://chat.dev/

Category
model serving
Stars
489
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 815 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence deep-learning distillation distillation-model llm llm-agent

hpcaitech/SwiftInfer

Efficient AI Inference & Serving

Category
model serving
Stars
478
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 942 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence deep-learning gpt inference llama llama2

ray-project/ray-educational-materials

This is suite of the hands-on training materials that shows how to scale CV, NLP, time-series forecasting workloads with Ray.

Category
model serving
Stars
460
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 907 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deep-learning distributed-machine-learning generative-ai llm llm-inference llm-serving

huawei-csl/KVarN

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

Category
model serving
Stars
452
Readiness
needs review (61/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +18 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai kv-cache llm llm-inference long-context quantization

AI-Hypercomputer/JetStream

JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).

Category
model serving
Stars
451
Readiness
needs review (56/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 214 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gemma gpt gpu inference jax large-language-models

FlagAI-Open/Aquila2

The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.

Category
model serving
Stars
447
Readiness
high risk (36/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load. Risks: no push in 665 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference llm-training

rizerphe/local-llm-function-calling

A tool for generating function arguments and choosing what function to call with local LLMs

Category
model serving
Stars
435
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 878 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

chatgpt-functions huggingface-transformers json-schema llm llm-inference openai-function-call

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

Category
model serving
Stars
398
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
17/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 103 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

delta-compression embeddable gguf kv-cache llm llm-inference

NVIDIA/Star-Attention

Efficient LLM Inference over Long Sequences

Category
model serving
Stars
391
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 408 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

attention-mechanism large-language-models llm-inference

DebarghaG/proofofthought

Proof of thought : LLM-based reasoning using Z3 theorem proving with multiple backend support (SMT2 and JSON DSL)

Category
model serving
Stars
376
Readiness
needs review (71/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automated-reasoning llm llm-inference llm-reasoning trustworthy-ai z3

alipay/PainlessInferenceAcceleration

Accelerate inference without tears

Category
model serving
Stars
371
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 196 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm-inference

morpheuslord/HackBot

AI-powered cybersecurity chatbot designed to provide helpful and accurate answers to your cybersecurity-related queries and also do code analysis and scan analysis.

Category
model serving
Stars
359
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
28/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 167 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai automation chatbot cli-chat-app cybersecurity cybersecurity-education

intel/neural-speed

An innovative library for efficient LLM inference via low-bit quantization

Category
model serving
Stars
352
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 707 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpu fp4 fp8 gaudi2 gpu int1

MLSys-Learner-Resources/Awesome-MLSys-Blogger

The repository has collected a batch of noteworthy MLSys bloggers (Algorithms/Systems)

Category
model serving
Stars
341
Readiness
high risk (38/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
24/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 579 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference llm-training machine-learning machine-learning-systems mlsys

structuredllm/syncode

Efficient and general syntactical decoding for Large Language Models

Category
model serving
Stars
338
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 200 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

grammar large-language-models llm llm-inference parser

interestingLSY/swiftLLM

A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).

Category
model serving
Stars
331
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 423 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda gpt inference inference-engine llama llm

andrewkchan/deepseek.cpp

CPU inference for the DeepSeek family of large language models in C++

Category
model serving
Stars
317
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 309 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpp deepseek llama llm llm-inference machine-learning

ByteDance-Seed/ShadowKV

[ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Category
model serving
Stars
311
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 463 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpu-offload high-throughput llm-inference long-context low-rank research

lambda-calculus-LLM/lambda-RLM

Method for Long Context RLMs using verifiable Lambda Calculus

Category
model serving
Stars
305
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
17/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 105 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

document-qa functional-programming lambda-calculus llm llm-inference long-context

galeselee/Awesome_LLM_System-PaperList

Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of papers on accelerating LLMs, currently focusing mainly on inferenc

Category
model serving
Stars
284
Readiness
high risk (40/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
21/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 519 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm-inference llm-serving paperlist papers system

Infini-AI-Lab/TriForce

[COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Category
model serving
Stars
281
Readiness
high risk (40/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 706 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

acceleration efficiency inference llm llm-inference long-context

modelscope/dash-infer

DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures, including CUDA, x86 and ARMv9.

Category
model serving
Stars
273
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 366 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpu cuda guided-decoding llm llm-inference native-engine

jax-ml/jax-llm-examples

Minimal yet performant LLM examples in pure JAX

Category
model serving
Stars
271
Readiness
needs review (67/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

jax llm llm-inference

efeslab/fiddler

[ICLR'25] Fast Inference of MoE Models with CPU-GPU Orchestration

Category
model serving
Stars
266
Readiness
needs review (56/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 628 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference local-inference mixtral-8x7b mixture-of-experts

ubergarm/r1-ktransformers-guide

run DeepSeek-R1 GGUFs on KTransformers

Category
model serving
Stars
258
Readiness
high risk (38/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, documentation. Risks: no push in 522 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deepseek-r1 ktransformers llm-inference llms

Infini-AI-Lab/MagicPIG

[ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation

Category
model serving
Stars
255
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 599 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

decoding gpu-cpu llm-inference lsh-algorithm

deeppowers/deeppowers

DEEPPOWERS is a Fully Homomorphic Encryption (FHE) framework built for MCP (Model Context Protocol), aiming to provide end-to-end privacy protection and high-efficiency computation for the upstream an

Category
model serving
Stars
252
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 469 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

accelerator ai llm-inference llms

inferflow/inferflow

Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).

Category
model serving
Stars
251
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 875 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

baichuan2 bloom deepseek falcon gemma internlm

Adriankhl/godot-llm

LLM in Godot

Category
model serving
Stars
251
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 775 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

game-development gamedev gdextension godot godot-engine godotengine

bytedance/ABQ-LLM

An acceleration library that supports arbitrary bit-width combinatorial quantization operations

Category
model serving
Stars
247
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 676 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda llm-inference mlsys quantized-networks research

numindai/nuextract

No description

Category
model serving
Stars
246
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +22 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

information-extraction llm llm-inference llm-training machine-learning nlp

little51/llm-dev

《大模型项目实战:多领域智能应用开发》配套资源

Category
model serving
Stars
233
Readiness
needs review (59/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
25/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 147 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

chat-application llm llm-deployment llm-inference llm-training

MorpheusAIs/Morpheus

Morpheus - A Network For Powering Smart Agents - Compute + Code + Capital + Community

Category
model serving
Stars
227
Readiness
needs review (60/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 778 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai compute ethereum llm-inference llms

devnen/qwen3.6-windows-server

One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.

Category
model serving
Stars
227
Readiness
high risk (41/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm-inference local-llm offline-ai privacy qwen qwen3

arc53/llm-price-compass

This project collects GPU benchmarks from various cloud providers and compares them to fixed per token costs. Use our tool for efficient LLM GPU selections and cost-effective AI models. LLM provider p

Category
model serving
Stars
223
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: no push in 599 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

benchmark gpu hacktoberfest inference-comparison llm llm-comparison

C0deMunk33/bespoke_automata

Bespoke Automata is a GUI and deployment pipline for making complex AI agents locally and offline

Category
model serving
Stars
222
Readiness
needs review (45/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
28/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation. Risks: no push in 166 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai automation chatbots developer-tools llm-inference

FasterDecoding/REST

REST: Retrieval-Based Speculative Decoding, NAACL 2024

Category
model serving
Stars
220
Readiness
needs review (47/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
26/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, license. Risks: no push in 155 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm-inference retrieval speculative-decoding

NimbleEdge/sparse_transformers

Sparse Inferencing for transformer based LLMs

Category
model serving
Stars
219
Readiness
needs review (49/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
23/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: no push in 135 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-inference sparsity transformers-models

dreadnode/burpference

A research project to add some brrrrrr to Burp

Category
model serving
Stars
212
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
29/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: no push in 172 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai burpsuite burpsuite-extension burpsuite-tools hackertools llm

Dynamis-Labs/spectralquant

SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression

Category
model serving
Stars
202
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

compression kv-cache large-language-models llm-inference machine-learning pytorch

aimhubio/aim

Aim 💫 — An easy-to-use & supercharged open-source experiment tracker.

Category
model serving
Stars
6,223
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai data-science data-visualization experiment-tracking machine-learning metadata

llama-farm/llamafarm

Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes

Category
model serving
Stars
835
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai aiproject chatgpt claude edge edge-computing

lance-format/lance

Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and

Category
model serving
Stars
6,919
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +28 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

apache-arrow computer-vision data-analysis data-analytics data-centric data-format

halfrost/Halfrost-Field

✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记

Category
model serving
Stars
13,215
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

algorithms blog cryptography go golang http2

LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Category
model serving
Stars
11,060
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +98 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

amd cuda fast inference kv-cache llm

OpenRLHF/OpenRLHF

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Category
model serving
Stars
9,894
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +27 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

large-language-models proximal-policy-optimization raylib reinforcement-learning reinforcement-learning-from-human-feedback transformers

xorbitsai/inference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready

Category
model serving
Stars
9,483
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence chatglm deployment flan-t5 gemma ggml

kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

Category
model serving
Stars
6,204
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +94 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

disaggregation inference kvcache llm rdma reinforcement-learning

mostlygeek/llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

Category
model serving
Stars
5,291
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +72 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

golang llama llamacpp localllama localllm openai

katanaml/sparrow

Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM

Category
model serving
Stars
5,192
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai computer-vision documentai huggingface-transformers llm machinelearning

skyzh/tiny-llm

A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.

Category
model serving
Stars
4,446
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +19 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

course large-language-model llm python qwen serving

PaddlePaddle/FastDeploy

High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

Category
model serving
Stars
3,702
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ernie ernie-45 ernie-45-vl inference llm llm-serving

containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of con

Category
model serving
Stars
2,990
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai containers cuda hacktoberfest hip inference-server

vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

Category
model serving
Stars
2,580
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +44 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ascend inference llm llm-serving llmops mlops

sybil-solutions/local-studio

Control panel for VLLM, Sglang, llama.cpp, exllamav3

Category
model serving
Stars
1,598
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +76 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai exllama hosting llamacpp local local-ai

intel/auto-round

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.

Category
model serving
Stars
1,557
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +12 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

diffusers gguf int4 llms mxfp4 nvfp4

SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 —

Category
model serving
Stars
1,341
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +36 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai amd benchmark cuda deepseek gb200

ModelCloud/GPTQModel

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Category
model serving
Stars
1,222
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gptq optimum peft quantization sglang transformers

Tencent-Hunyuan/UniRL

UniRL is a Framework for Unified Multimodal Model Reinforcement Learning

Category
model serving
Stars
889
Readiness
ready (78/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +21 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-infrastructure reinforcement-learning sglang vllm

verl-project/verl-omni

Multimodal RL training framework for diffusion & omni models

Category
model serving
Stars
748
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +57 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

diffusion-models flow-matching grpo multimodal qwen reinforcement-learning

raids-lab/crater

Crater is a cloud-native AI training & inference platform.

Category
model serving
Stars
546
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

buildkit deep-learning envd fumadocs helm-charts jupyter-notebook

runpod-workers/worker-vllm

The Runpod worker template for serving our large language model endpoints. Powered by vLLM.

Category
model serving
Stars
461
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
8/100 · high confidence

Why: +1 stars in 7 days; 23 commits in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity; open issue backlog is stable or shrinking

Strongest signals: push recency, commit activity, contributor breadth. Risks: no pull-request review responses in 30 days. Missing inputs: None.

language-model llm runpod vllm

spark-arena/sparkrun

sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems

Category
model serving
Stars
431
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
8/100 · high confidence

Why: +16 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity

Strongest signals: push recency, commit activity, contributor breadth. Risks: no pull-request review responses in 30 days. Missing inputs: None.

Capped lower bounds: 30-day commits.

dgx-spark inference llama-cpp sglang vllm

guqiong96/Lvllm

LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, suppo

Category
model serving
Stars
428
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +38 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: 30-day commits, lifetime contributors.

cpu decode gpu hybrid inference model

FujitsuResearch/OneCompression

Python package for LLM compression

Category
model serving
Stars
415
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +17 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm qep quantization vllm

ModelEngine-Group/unified-cache-management

Persist and reuse KV Cache to speedup your LLM.

Category
model serving
Stars
313
Readiness
ready (96/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ascend cuda deepseek dram gpu hbm

pensarai/apex

AI-powered offensive security testing using autonomous agents, directly in your terminal.

Category
model serving
Stars
300
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai ai-sdk anthropic cybersecurity offensive-security

guoqingbao/xinfer

Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.

Category
model serving
Stars
299
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent llm qwen rust vllm

mixa3607/ML-gfx906

ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60

Category
model serving
Stars
297
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

amd-gpu comfyui llamacpp torch vllm

vortico/flama

The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.

Category
model serving
Stars
296
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

anthropic asgi chatbot domain-driven-design generative-ai inference

lightseekorg/TorchSpec

A PyTorch native library for training speculative decoding models

Category
model serving
Stars
223
Readiness
ready (100/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

dspark eagle3 fsdp lightseek llm mooncake

Roots-Automation/GutenOCR

Open-source tools for training and evaluating Vision Language Models for OCR

Category
model serving
Stars
190
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llms multigpu ocr vllm vlm-ocr vlms

novitalabs/pegaflow

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

Category
model serving
Stars
184
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

inference kv-cache llm vllm

OpenDCAI/One-Eval

Automated system for LLM evaluation via agents. Doc as below:

Category
model serving
Stars
160
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +10 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agents benchmark data data-analysis data-science

Sandermage/sndr_core_engine

SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B

Category
model serving
Stars
131
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awq consumer-gpu cuda gemma kv-cache-quantization llm-inference

pmady/keda-gpu-scaler

KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required

Category
model serving
Stars
113
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-infrastructure autoscaling daemonset gpu gpu-autoscaling gpu-metrics

intel/optimization-zone

Data Center and Client workload and software optimizations for Intel hardware.

Category
model serving
Stars
107
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

benchmarking cassandra envoy intel java kafka

VectorInstitute/vector-inference

Efficient LLM inference on Slurm clusters.

Category
model serving
Stars
106
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

audio-transcription inference llm llm-infernece llm-infrastructure multimodal

ai-runway/airunway

✈️ Kubernetes-native platform for deploying and managing AI inference across multiple providers

Category
model serving
Stars
96
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai cloudnative dynamo inference kubernetes llm

aws-samples/sample-genai-on-eks-starter-kit

A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: 🚀 AI Gateway (LiteLLM) 🤖 LLM Serving (vLLM, SGLang, Ollama)

Category
model serving
Stars
93
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai ai-agents ai-engineering ai-gateway ai-platform amazon-eks

generative-computing/granite-switch

Granite Switch — Build AI models like you build software

Category
model serving
Stars
90
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

generative-computing llm-inference llms transformers vllm

niklasfrick/spark-dashboard

Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.

Category
model serving
Stars
88
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-monitoring dashboard dgx gpu gpu-monitoring

helasaoudi/llm-inspector

The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.

Category
model serving
Stars
66
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gpu-monitoring htop inference llm llm-inference llminspect

Orchestra-Research/AI-Research-SKILLs

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower.

Category
model serving
Stars
11,495
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +204 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-research claude claude-code claude-skills codex

th1nhhdk/local_ai_ocr

An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).

Category
model serving
Stars
779
Readiness
ready (79/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai deepseek-ocr deepseek-ocr-2 docx english llm

ModelTC/LightCompress

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

Category
model serving
Stars
739
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awq benchmark deepseek-v3 deployment evaluation internlm2

microsoft/vidur

Accurate, large-scale, and extensible simulator for LLM inference Systems

Category
model serving
Stars
657
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 378 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

inference llm simulation transformer vllm

HuiResearch/FlashTTS

基于SparkTTS、OrpheusTTS等模型,提供高质量中文语音合成与声音克隆服务。

Category
model serving
Stars
609
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, documentation. Risks: no push in 446 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

flashtts llamacpp-python megatts3 orpheus-tts sglang spark-tts

FlagOpen/RoboBrain

[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository.

Category
model serving
Stars
560
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 298 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

embodied-ai robotics vllm

micytao/vllm-playground

A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterp

Category
model serving
Stars
507
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
20/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 122 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai learning llms vllm

neosun100/DeepSeek-OCR-WebUI

🎨 Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | Real-time Logging | Fully Responsive

Category
model serving
Stars
441
Readiness
needs review (59/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
28/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 168 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

batch-processing computer-vision deepseek fastapi image-recognition modern-ui

AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash

Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-co

Category
model serving
Stars
438
Readiness
needs review (70/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

abliteration blackwell dflash dgx-spark llm nvfp4

chtmp223/topicGPT

TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24]

Category
model serving
Stars
413
Readiness
needs review (61/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm nlp openai python topic-modeling vllm

varunshenoy/super-json-mode

Low latency JSON generation using LLMs ⚡️

Category
model serving
Stars
396
Readiness
high risk (36/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 880 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

huggingface-transformers llm openai vllm

yaof20/Flash-RL

Implementation for FP8/INT8 Rollout for RL training without performence drop.

Category
model serving
Stars
308
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 273 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

reinforcement-learning vllm

albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4

Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)

Category
model serving
Stars
306
Readiness
needs review (56/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

autoround cuda dgx-spark lossless mtp performance-optimization

SunzeY/SEAgent

[ICML-2026] Official implementation of "SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience"

Category
model serving
Stars
260
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 365 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent computer-use-agent grpo gui-agent osworld rl

lucasjinreal/Namo-R1

A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease.

Category
model serving
Stars
256
Readiness
high risk (41/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 472 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm mllm moondream smolvlm vllm vllms

JackYFL/awesome-VLLMs

This repository collects papers on VLLM applications. We will update new papers irregularly.

Category
model serving
Stars
221
Readiness
high risk (41/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

application embodied llm mllm reasoning-agent survey

ovshake/nano-vllm

a fun and educational take on vLLM

Category
model serving
Stars
215
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, license. Risks: no push in 194 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

inference-engine python vllm

kossisoroyce/timber

Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster t

Category
model serving
Stars
688
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
19/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 113 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

c99 catboost compiler decision-trees gradient-boosting inference

stas00/ml-engineering

Machine Learning Engineering Open Book

Category
model serving
Stars
18,530
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +29 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai debugging gpus inference large-language-models llm

SwanHubX/SwanLab

⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / Ultr

Category
model serving
Stars
4,125
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +22 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-infra data-science deep-learning llm logging machine-learning

huggingface/optimum

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

Category
model serving
Stars
3,455
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

graphcore habana inference intel onnx onnxruntime

openvinotoolkit/nncf

Neural Network Compression Framework for enhanced OpenVINO™ inference

Category
model serving
Stars
1,187
Readiness
ready (99/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bert classification compression deep-learning genai llm

SkalskiP/courses

This repository is a curated collection of links to various courses and resources about Artificial Intelligence (AI)

Category
model serving
Stars
6,476
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +13 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 837 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

computer-vision deep-learning deep-neural-networks generative-model machine-learning mlops

AutoGPTQ/AutoGPTQ

An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.

Category
model serving
Stars
5,075
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 483 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deep-learning inference large-language-models llms nlp pytorch

IntelLabs/nlp-architect

A model library for exploring state-of-the-art deep learning topologies and techniques for optimizing Natural Language Processing neural networks

Category
model serving
Stars
2,930
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 1369 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bert deep-learning deeplearning dynet nlp nlu

tjake/Jlama

Jlama is a modern LLM inference engine for Java

Category
model serving
Stars
1,298
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 299 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai genai gpt huggingface java llama

ModelTC/LightLLM

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

Category
model serving
Stars
4,213
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deep-learning gpt llama llm model-serving nlp

thu-pacman/chitu

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

Category
model serving
Stars
3,147
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +21 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deepseek gpu llm llm-serving model-serving pytorch

openlake-project/openlake

OpenLake is a high performance storage engine for efficient LLM inference and GPU Training

Category
model serving
Stars
2,305
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +533 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

blackwell gpt gpu high-performance llm llm-training

tensorchord/envd

🏕️ Reproducible development environment for humans and agents

Category
model serving
Stars
2,219
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent buildkit code-agent codex developer-tools development-environment

mlrun/mlrun

MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates t

Category
model serving
Stars
1,690
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-engineering data-science experiment-tracking kubernetes machine-learning mlops

kitops-ml/kitops

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

Category
model serving
Stars
1,395
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai code datasets devops devops-tools gguf

alibaba/rtp-llm

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

Category
model serving
Stars
1,297
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gpt inference llama llm llm-serving llmops

basetenlabs/truss

The simplest way to serve AI/ML models in production

Category
model serving
Stars
1,186
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence easy-to-use falcon inference-api inference-server machine-learning

openvinotoolkit/model_server

A scalable inference server for models optimized with OpenVINO™

Category
model serving
Stars
908
Readiness
ready (97/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai cloud dag deep-learning edge genai

mosecorg/mosec

A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

Category
model serving
Stars
903
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cv deep-learning gpu hacktoberfest jax llm

Tejas-TA/predikit

The missing bridge between your ML models and your AI agents.

Category
model serving
Stars
419
Readiness
ready (96/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents langchain llm machine-learning model-serving openai

alibaba/ServeGen

A framework for generating realistic LLM serving workloads

Category
model serving
Stars
168
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deepseek llm llm-serving model-serving qwen

jdaln/dgx-spark-inference-stack

Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for now and single-spark. For the not-so-rich buddies. If you want l

Category
model serving
Stars
51
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda dgx dgx-spark docker docker-compose gb10

snapllm/snapllm

🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffusion.cpp

Category
model serving
Stars
39
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai cpu-inference edge-ai llm llm-inference llm-serving

aivrar/multi-turboquant

Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi

Category
model serving
Stars
25
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

attention compression cuda deep-learning gpu inference

togettoyou/kpilot

KPilot: Unified control plane for multi-cluster Kubernetes management, GPU compute scheduling, and model serving.

Category
model serving
Stars
25
Readiness
ready (74/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent ai batch-systems cloud-native gpu-management gpu-shareable

paralleliq/piqc

Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.

Category
model serving
Stars
23
Readiness
ready (73/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-infrastructure cloud-native gpu introspection kubernetes llm

batchgen-project/batchgen

High-Throughput Batch Inference

Category
model serving
Stars
13
Readiness
needs review (71/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deepseek deepseek-r1 deepseek-v3 large-language-models mixture-of-experts model-serving

joeynyc/honeycomb-lab

Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet

Category
model serving
Stars
13
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

dgx-spark gpu homelab llm-gateway lm-studio local-inference

noctaya/noctaya

Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs

Category
model serving
Stars
11
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ascend autoscaling cloud-native inference keda kubernetes

ai-lab-tech/triton-control

Web UI for NVIDIA Triton Inference Server on Kubernetes

Category
model serving
Stars
9
Readiness
ready (100/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai control-plane dashboard inference-server kubernetes mlops

Raytorin/triton-openai-gateway

OpenAI-compatible gateway for NVIDIA Triton and vLLM with tools, multimodal inputs, embeddings, reranking, and observability.

Category
model serving
Stars
8
Readiness
needs review (64/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

embeddings fastapi gpu-inference kubernetes llm-gateway llm-serving

Mizistein/omlx

🤖 Optimize LLM inference on Mac with continuous batching and SSD caching managed from your menu bar for efficient performance.

Category
model serving
Stars
8
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

apple-silicon chatbot inference-server llm macos mlx

chaeminyoon/AIS-Traffic-Ops

An MLOps project for forecasting maritime traffic patterns from AIS data and operating the model through evaluation, serving, and monitoring workflows.

Category
model serving
Stars
5
Readiness
needs review (61/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ais fastapi grafana maritime mlflow mlops

open-nvr/ai-adapter

AI Adapter is a flexible integration layer for connecting any AI model to OpenNVR. It enables seamless support for cloud, local, and edge models through a modular architecture—allowing developers to p

Category
model serving
Stars
5
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai computer-vision docker face-recognition inference inference-server

chriss1245/mlflow_packaging_samples

Worked examples of packaging PyTorch models with MLflow — native flavour vs custom pyfunc, and what each costs at serving time.

Category
model serving
Stars
5
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

machine-learning mlflow mlops model-packaging model-serving pytorch

ahkarami/Deep-Learning-in-Production

In this repository, I will share some useful notes and references about deploying deep learning-based models in production.

Category
model serving
Stars
4,375
Readiness
needs review (45/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 636 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

angularjs c-plus-plus caffe2 convert-pytorch-models deep-learning deep-neural-networks

logicalclocks/hopsworks

Hopsworks - Data-Intensive AI platform with a Feature Store

Category
model serving
Stars
1,301
Readiness
needs review (47/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 543 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aws azure data-science feature-engineering feature-management feature-store

efeslab/Nanoflow

A throughput-oriented high-performance serving framework for LLMs

Category
model serving
Stars
974
Readiness
high risk (39/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
22/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 131 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda inference llama2 llm llm-serving model-serving

bentoml/Yatai

Model Deployment at Scale on Kubernetes 🦄️

Category
model serving
Stars
841
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
70/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, documentation. Risks: repository is archived. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bentoml k8s kubernetes machine-learning mlops model-deployment

ServerlessLLM/ServerlessLLM

Serverless LLM Serving for Everyone.

Category
model serving
Stars
700
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
16/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 95 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda huggingface-transformers large-language-models model-as-a-service model-serving pytorch

eightBEC/fastapi-ml-skeleton

FastAPI Skeleton App to serve machine learning models production-ready.

Category
model serving
Stars
604
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 211 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

fastapi machine-learning model-serving python python3

underneathall/pinferencia

Python + Inference - Model Deployment library in Python. Simplest model inference server ever.

Category
model serving
Stars
543
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 1270 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai artificial-intelligence computer-vision data-science deep-learning huggingface

intel/xFasterTransformer

No description

Category
model serving
Stars
435
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 323 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

chatglm inference intel llama llm model-serving

Lightning-Universe/stable-diffusion-deploy

Learn to serve Stable Diffusion models on cloud infrastructure at scale. This Lightning App shows load-balancing, orchestrating, pre-provisioning, dynamic batching, GPU-inference, micro-services worki

Category
model serving
Stars
391
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: repository is archived, no push in 1039 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

model-serving stable-diffusion

aniketmaurya/chitra

A multi-functional library for full-stack Deep Learning. Simplifies Model Building, API development, and Model Deployment.

Category
model serving
Stars
235
Readiness
ready (75/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bounding-boxes deep-learning fastapi gradcam hacktoberfest image-classification

lightbend/kafka-with-akka-streams-kafka-streams-tutorial

Code samples for the Lightbend tutorial on writing microservices with Akka Streams, Kafka Streams, and Kafka

Category
model serving
Stars
208
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, fork interest, license. Risks: no push in 2627 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

akka kafka-streams model-serving

HumanSignal/label-studio

Label Studio is a multi-type data labeling and annotation tool with standardized output format

Category
model serving
Stars
28,009
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +45 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

annotation annotation-tool annotations boundingbox computer-vision data-labeling

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Category
model serving
Stars
20,831
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +23 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awesome awesome-list data-mining deep-learning explainability interpretability

microsoft/agent-lightning

The absolute trainer to light up AI agents.

Category
model serving
Stars
17,457
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +27 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agentic-ai llm mlops reinforcement-learning

DataTalksClub/mlops-zoomcamp

Free MLOps course from DataTalks.Club

Category
model serving
Stars
15,079
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +36 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

machine-learning mlops model-deployment model-monitoring workflow-orchestration

wandb/wandb

The AI developer platform. Use Weights & Biases to train and fine-tune models, and manage models from experimentation to production.

Category
model serving
Stars
11,219
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai collaboration data-science data-versioning deep-learning experiment-track

aws/amazon-sagemaker-examples

Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.

Category
model serving
Stars
10,981
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aws data-science deep-learning examples inference jupyter-notebook

kedro-org/kedro

Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, an

Category
model serving
Stars
10,950
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +10 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai agentic-workflow data-pipelines hacktoberfest kedro machine-learning

skypilot-org/skypilot

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

Category
model serving
Stars
10,461
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +33 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cloud-computing cloud-management cost-optimization deep-learning distributed-training gpu

pycaret/pycaret

Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.

Category
model serving
Stars
9,833
Readiness
ready (78/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

anomaly-detection automl classification clustering data-science fastapi

clearml/clearml

ClearML - Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution

Category
model serving
Stars
6,815
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai clearml control deep-learning deeplearning devops

zenml-io/zenml

ZenML 🙏: One AI Platform from Pipelines to Agents. https://zenml.io.

Category
model serving
Stars
5,545
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +17 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentops agents ai automl data-science deep-learning

tencentmusic/cube-studio

cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持

Category
model serving
Stars
5,076
Readiness
needs review (70/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +10 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai aihub argo automl deepseek gpt

kubeflow/pipelines

Machine Learning Pipelines for Kubeflow

Category
model serving
Stars
4,181
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-science kubeflow kubeflow-pipelines kubernetes machine-learning mlops

polyaxon/polyaxon

AI Infra / AI Orchestration / AI Control Plane

Category
model serving
Stars
3,717
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents artificial-intelligence data-science deep-learning harness hyperparameter-optimization

datachain-ai/datachain

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Category
model serving
Stars
2,805
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents claude-code codex data-context-layer data-processing harness-engineering

apache/burr

Build applications that make decisions (chatbots, agents, simulations, etc...). Monitor, trace, persist, and execute on your own infrastructure.

Category
model serving
Stars
2,504
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai burr chatbot-framework dags generative-ai graphs

akuity/awesome-argo

A curated list of awesome projects and resources related to Argo (a CNCF graduated project)

Category
model serving
Stars
2,467
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

argo argo-events argo-rollouts argo-workflows argocd awesome

data-infra/cube-studio

cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模

Category
model serving
Stars
2,400
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-platform ascend automl cube-studio cubestudio deepseek

lupinemachines/lupine

LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

Category
model serving
Stars
2,373
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cublas cuda cudnn gpu mlops networking

kubeflow/katib

Automated Machine Learning on Kubernetes

Category
model serving
Stars
1,694
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai automl huggingface hyperparameter-tuning jax kubeflow

psalias2006/gpu-hot

🔥 Real-time NVIDIA GPU dashboard

Category
model serving
Stars
1,593
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

charts cuda dashboard devops docker flask

MLSysOps/MLE-agent

🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 OpenAI, Anthropic, Gemini, Ollama,

Category
model serving
Stars
1,565
Readiness
ready (72/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent ai llm ml mle mlops

AgileRL/AgileRL

Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimization.

Category
model serving
Stars
942
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents agilerl automl deep-learning deep-reinforcement-learning distributed

lightly-ai/lightly-studio

Curate, Annotate, and Manage Your Data in LightlyStudio.

Category
model serving
Stars
874
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

computer-vision image-labeling mlops

GoogleCloudPlatform/vertex-ai-samples

Notebooks, code samples, sample apps, and other resources that demonstrate how to use, develop and manage machine learning and generative AI workflows using Google Cloud Vertex AI.

Category
model serving
Stars
774
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automl colab colab-enterprise gemini gemini-api genai

nuclia/nucliadb

NucliaDB, The AI Search database for RAG

Category
model serving
Stars
719
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-powered-search database language-model machine-learning mlops nuclia

statmike/vertex-ai-mlops

Google Cloud Platform Vertex AI end-to-end workflows for machine learning operations

Category
model serving
Stars
709
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deep-learning gcp gcp-vertex-ai machine-learning mlops mlops-template

databricks/mlops-stacks

This repo provides a customizable stack for starting new ML projects on Databricks that follow production best-practices out of the box.

Category
model serving
Stars
708
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

databricks machine-learning mlops

Meesho/BharatMLStack

BharatMLStack is an open-source, end-to-end machine learning infrastructure stack built at Meesho to support real-time and batch ML workloads at Bharat scale

Category
model serving
Stars
705
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai feature-engineering feature-store-online machine-learning ml mlops

polyaxon/traceml

Engine for AI/ML/Data tracking, visualization, explainability, drift detection, and dashboards for Polyaxon.

Category
model serving
Stars
534
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

dask data-exploration data-profiling data-quality data-quality-checks data-science

CerebriumAI/examples

Examples for Cerebrium Serverless GPUs

Category
model serving
Stars
526
Readiness
needs review (68/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai gpu llms ml mlops serverless

skops-dev/skops

skops is a Python library helping you share your scikit-learn based models and put them in production

Category
model serving
Stars
524
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

hacktoberfest huggingface machine-learning mlops scikit-learn

techiescamp/mlops-for-devops

MLOps for DevOps Engineers - A hands-on, project-based guide to Machine Learning Operations

Category
model serving
Stars
508
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +13 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

devops devops-mlops mlops mlops-project

operatorai/modelstore

🏬 modelstore is a Python library that allows you to version, export, and save a machine learning model to your filesystem or a cloud storage provider.

Category
model serving
Stars
404
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-science keras machine-learning mlops modelstore python-library

rodrigo-arenas/Sklearn-genetic-opt

Hyperparameter optimization and feature selection for scikit-learn using evolutionary algorithms. A modern alternative to GridSearchCV and RandomizedSearchCV.

Category
model serving
Stars
386
Readiness
ready (98/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence automl cross-validation evolutionary-algorithms feature-engineering feature-selection

m3dev/gokart

Gokart solves reproducibility, task dependencies, constraints of good code, and ease of use for Machine Learning Pipeline.

Category
model serving
Stars
342
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gokart luigi machine-learning mlops pipeline-framework

flyteorg/flytekit

Extensible Python SDK for developing Flyte tasks and workflows. Simple to get started and learn and highly extensible.

Category
model serving
Stars
314
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automation data data-science extensible flyte flyte-tasks

hongbo-miao/hongbomiao.com

A personal research and development (R&D) lab that facilitates the sharing of knowledge.

Category
model serving
Stars
298
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aerospace autonomy cloud-native computational-fluid-dynamics computer-vision distributed-tracing

microsoft/nni

An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.

Category
model serving
Stars
14,364
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 765 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automated-machine-learning automl bayesian-optimization data-science deep-learning deep-neural-network

visenger/awesome-mlops

A curated list of references for MLOps

Category
model serving
Stars
14,135
Readiness
needs review (45/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
26/100 · low confidence

Why: +21 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 624 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai data-science devops engineering federated-learning machine-learning

chiphuyen/machine-learning-systems-design

A booklet on machine learning systems design with exercises. NOT the repo for the book "Designing Machine Learning Systems", which is `dmls-book`

Category
model serving
Stars
10,485
Readiness
high risk (40/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +12 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load. Risks: no push in 1210 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-science machine-learning-production mlops

tensorchord/Awesome-LLMOps

An awesome & curated list of best LLMOps tools for developers

Category
model serving
Stars
5,910
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-development-tools awesome-list llmops mlops

ashleve/lightning-hydra-template

PyTorch Lightning + Hydra. A very user-friendly template for ML experimentation. ⚡🔥⚡

Category
model serving
Stars
5,329
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

best-practices config deep-learning hydra mlops project-structure

kelvins/awesome-mlops

:sunglasses: A curated list of awesome MLOps tools

Category
model serving
Stars
5,230
Readiness
high risk (44/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai awesome data-science machine-learning machine-learning-engineering ml

SeldonIO/seldon-core

An MLOps framework to package, deploy, monitor and manage thousands of production machine learning models

Category
model serving
Stars
4,767
Readiness
needs review (49/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
23/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 137 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aiops deployment kubernetes machine-learning machine-learning-operations mlops

pytorch/serve

Serve, optimize and scale PyTorch models in production

Category
model serving
Stars
4,349
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 366 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cpu deep-learning docker gpu kubernetes machine-learning

deepchecks/deepchecks

Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and model

Category
model serving
Stars
4,040
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: +7 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-drift data-science data-validation deep-learning html-report jupyter-notebook

higgsfield-ai/higgsfield

Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters

Category
model serving
Stars
4,036
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +29 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 804 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cluster-management deep-learning distributed llama llama2 llm

determined-ai/determined

Determined is an open-source machine learning platform that simplifies distributed training, hyperparameter tuning, experiment tracking, and resource management. Works with PyTorch and TensorFlow.

Category
model serving
Stars
3,231
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 505 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-science deep-learning distributed-training hyperparameter-optimization hyperparameter-search hyperparameter-tuning

plexe-ai/plexe

✨ Build a machine learning model from a prompt

Category
model serving
Stars
2,593
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
26/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 154 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai agents ai machine-learning ml mlengineering

NannyML/nannyml

nannyml: post-deployment data science in python

Category
model serving
Stars
2,147
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
22/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 391 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-analysis data-drift data-science deep-learning jupyter-notebook machine-learning

microsoft/MLOps

MLOps examples

Category
model serving
Stars
2,112
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +6 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, license. Risks: no push in 735 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

azureml mlops

premAI-io/state-of-open-source-ai

:closed_book: Clarity in the current fast-paced mess of Open Source innovation

Category
model serving
Stars
1,636
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 564 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai book hacktoberfest jupyter-book ml mlops

ai-infra-curriculum/ai-infra-engineer-learning

AI Infrastructure Engineer Learning Track - Production ML infrastructure curriculum (2-4 years experience)

Category
model serving
Stars
1,559
Readiness
needs review (69/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +25 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-infrastructure career-development curriculum devops education

modelfoxdotdev/modelfox

ModelFox makes it easy to train, deploy, and monitor machine learning models.

Category
model serving
Stars
1,466
Readiness
high risk (42/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 735 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automl developer-tools elixir elixir-lang go golang

MLReef/mlreef

The collaboration workspace for Machine Learning

Category
model serving
Stars
1,458
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 1375 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence data-science deep-learning deeplearning machine-learning machine-learning-algorithms

ebhy/budgetml

Deploy a ML inference service on a budget in less than 10 lines of code.

Category
model serving
Stars
1,343
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 907 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

api data-science deployment fastapi inference machine-learning

microsoft/MLOpsPython

MLOps using Azure ML Services and Azure DevOps

Category
model serving
Stars
1,318
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +3 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, license. Risks: no push in 1098 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

azure-machine-learning azureml mlops

marsupialtail/quokka

Making data lake work for time series

Category
model serving
Stars
1,192
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 716 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-lake-analytics distributed etl-framework mlops sql

mryab/efficient-dl-systems

Efficient Deep Learning Systems course materials

Category
model serving
Stars
1,020
Readiness
needs review (57/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cuda deep-learning distributed-training efficient-deep-learning inference-optimization machine-learning

dillionverma/llm.report

📊 llm.report is an open-source logging and analytics platform for OpenAI: Log your ChatGPT API requests, analyze costs, and improve your prompts.

Category
model serving
Stars
1,020
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: repository is archived, no push in 815 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aiops gpt-3 gpt-4 llm llmops mlops

getmetal/motorhead

🧠 Motorhead is a memory and information retrieval server for LLMs.

Category
model serving
Stars
916
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 381 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llmops llms machine-learning ml mlops rust

Paulescu/hands-on-train-and-deploy-ml

Train and Deploy an ML REST API to predict crypto prices, in 10 steps

Category
model serving
Stars
888
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 800 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

crypto deployment ml mlops

zszazi/Deep-learning-in-cloud

List of Deep Learning Cloud Providers

Category
model serving
Stars
819
Readiness
needs review (64/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artificial-intelligence cloud cloud-gpus deep-learning deeplearning gpu

MLOps-Courses/mlops-coding-course

Learn how to create, develop, and maintain a state-of-the-art MLOps code base

Category
model serving
Stars
733
Readiness
needs review (68/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

best-practices coding courses data-science machine-learning mkdocs

zenml-io/awesome-open-data-annotation

Open Source Data Annotation & Labeling Tools

Category
model serving
Stars
719
Readiness
needs review (71/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +6 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai annotation datacentric labelled-data labelling machine-learning

cresset-template/cresset

Template repository to build PyTorch projects from source on any version of PyTorch/CUDA/cuDNN.

Category
model serving
Stars
718
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 573 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

build cuda deep-learning deep-learning-tutorial docker docker-compose

arthurhenrique/cookiecutter-fastapi

Cookiecutter template for FastAPI projects using: Machine Learning, uv, Github Actions and Pytests

Category
model serving
Stars
709
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 347 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai black boilerplate cli cookiecutter cookiecutter-fastapi

farukalamai/advanced-machine-learning-engineer-roadmap-2024

A Full Stack ML (Machine Learning) Roadmap involves learning the necessary skills and technologies to become proficient in all aspects of machine learning, including data collection and preprocessing,

Category
model serving
Stars
700
Readiness
needs review (56/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 704 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aws computer-vision data-analysis data-science data-visualization deep-learning

featurestoreorg/serverless-ml-course

Serverless Machine Learning Course for building AI-enabled Prediction Services from models and features

Category
model serving
Stars
685
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 682 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

course feature-engineering feature-store machine-learning ml mlops

Azure/mlops-v2

Azure MLOps (v2) solution accelerators. Enterprise ready templates to deploy your machine learning models on the Azure Platform.

Category
model serving
Stars
649
Readiness
needs review (61/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

azure azuremachinelearning azureml deep-learning devops machine-learning

cdfoundation/sig-mlops

CDF SIG MLOps

Category
model serving
Stars
633
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 615 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cdf cicd devop machine-learning ml mlops

smazzanti/mrmr

mRMR (minimum-Redundancy-Maximum-Relevance) for automatic feature selection at scale.

Category
model serving
Stars
630
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 626 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-science feature-selection machine-learning mlops

neptune-ai/neptune-client

📘 The experiment tracker for foundation model training

Category
model serving
Stars
623
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
94/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 143 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

comparison dl foundation keras learning lightgbm

Paulescu/kubernetes-for-ml-engineers

Just enough Kubernetes for you to fly

Category
model serving
Stars
579
Readiness
needs review (47/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, fork interest, documentation. Risks: no push in 497 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

docker fastapi kubernetes ml mlops

jacopotagliabue/MLSys-NYU-2022

Slides, scripts and materials for the Machine Learning in Finance Course at NYU Tandon, 2022

Category
model serving
Stars
558
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 1335 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

coursework fraud-detection introduction-to-machine-learning machine-learning mlops recommender-system

the-full-stack/fsdl-text-recognizer-2022-labs

Complete deep learning project developed in Full Stack Deep Learning, 2022 edition. Generated automatically from https://github.com/full-stack-deep-learning/fsdl-text-recognizer-2022

Category
model serving
Stars
530
Readiness
needs review (58/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, license. Risks: no push in 933 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

deep-neural-networks machine-learning mlops

RunLLM/aqueduct

Aqueduct is no longer being maintained. Aqueduct allows you to run LLM and ML workloads on any cloud infrastructure.

Category
model serving
Stars
517
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 1157 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai data data-science kubernetes llm llms

terrytangyuan/distributed-ml-patterns

Distributed Machine Learning Patterns from Manning Publications by Yuan Tang https://bit.ly/2RKv8Zo

Category
model serving
Stars
513
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

argo argo-workflows book cloud-computing cloud-native data-science

ing-bank/popmon

Monitor the stability of a Pandas or Spark dataframe ⚙︎

Category
model serving
Stars
512
Readiness
needs review (52/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 210 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

covariate-shift data-analysis data-distributions data-profiling data-science dataset-shifts

fuzzylabs/awesome-open-mlops

The Fuzzy Labs guide to the universe of open source MLOps

Category
model serving
Stars
482
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
18/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 445 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

datascience devops infrastructure machine-learning machinelearning mlops

zzsza/Boostcamp-AI-Tech-Product-Serving

부스트캠프 AI Tech - Product Serving 자료

Category
model serving
Stars
474
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: issue load, fork interest. Risks: no push in 225 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

mlops serving

loglabs/mltrace

Coarse-grained lineage and tracing for machine learning pipelines.

Category
model serving
Stars
469
Readiness
needs review (51/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 1365 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

devops machine-learning mlops pipeline-management tracing

deployKF/deployKF

deployKF builds machine learning platforms on Kubernetes. We combine the best of Kubeflow, Airflow†, and MLflow† into a complete platform.

Category
model serving
Stars
466
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 735 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

argocd artificial-intelligence gitops kubeflow kubernetes machine-learning

databrickslabs/dbx

🧱 Databricks CLI eXtensions - aka dbx is a CLI tool for development and advanced Databricks workflows management.

Category
model serving
Stars
462
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
22/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 133 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ci cicd databricks databricks-api databricks-cli mlops

encord-team/encord-active

The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.

Category
model serving
Stars
460
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 441 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

active-learning annotations computer-vision data data-centric data-cleaning

Writesonic/GPTRouter

Smoothly Manage Multiple LLMs (OpenAI, Anthropic, Azure) and Image Models (Dall-E, SDXL), Speed Up Responses, and Ensure Non-Stop Reliability.

Category
model serving
Stars
455
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 849 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

anthropic azure-openai cohere google-gemini langchain llama-index

polyaxon/haupt

Lineage metadata API, artifacts streams, sandbox, API, and spaces for Polyaxon

Category
model serving
Stars
452
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

bokeh data-processing data-profiling data-science data-visualization deep-learning

aporia-ai/mlplatform-workshop

🍫 Example code for a basic ML Platform based on Pulumi, FastAPI, DVC, MLFlow and more

Category
model serving
Stars
444
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 1738 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

devops machine-learning mlops

fmind/cookiecutter-mlops-package

Start building and deploying Python packages and Docker images for MLOps tasks.

Category
model serving
Stars
441
Readiness
ready (74/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

best-practices cookiecutter machine-learning mlflow mlops python

bodywork-ml/bodywork-core

ML pipeline orchestration and model deployments on Kubernetes.

Category
model serving
Stars
436
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: repository is archived, no push in 1085 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

batch cicd continuous-deployment data-science devops framework

noahgift/Python-MLOps-Cookbook

This is an example of a Containerized Flask Application that can deploy to many target environments including: AWS, GCP and Azure.

Category
model serving
Stars
435
Readiness
needs review (47/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest. Risks: no push in 574 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cookbook mlops

arthur-ai/bench

A tool for evaluating LLMs

Category
model serving
Stars
428
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
24/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 145 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm mlops

GoogleCloudPlatform/mlops-with-vertex-ai

An end-to-end example of MLOps on Google Cloud using TensorFlow, TFX, and Vertex AI

Category
model serving
Stars
424
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 828 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

gcp google-cloud-platform mlops tensorflow tfx vertex-ai

ashishpatel26/ResourceBank_CV_NLP_MLOPS_2022

This repository offers a goldmine of materials for students of computer vision, natural language processing, and machine learning operations.

Category
model serving
Stars
423
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 1376 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

computer-vision data-science deep-learning mlops natural-language-processing

AlexIoannides/kubernetes-mlops

MLOps tutorial using Python, Docker and Kubernetes.

Category
model serving
Stars
416
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 658 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

cloud-platform docker flask gcp helm kubernetes

kennethleungty/MLOps-Specialization-Notes

Notes for Machine Learning Engineering for Production (MLOps) Specialization course by DeepLearning.AI & Andrew Ng

Category
model serving
Stars
405
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 1178 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

andrew-ng course coursera data-science deep-learning deeplearningai

WarrenWen666/AI-Software-Startups

A Survey of AI startups

Category
model serving
Stars
400
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 1076 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai artificial-intelligence machine-learning mlops startup

thoughtworks/mlops-platforms

Compare MLOps Platforms. Breakdowns of SageMaker, VertexAI, AzureML, Dataiku, Databricks, h2o, kubeflow, mlflow...

Category
model serving
Stars
395
Readiness
high risk (0/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
100/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 1366 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

azureml data-science databricks dataiku datarobot google-ai-platform

mlop-ai/mlop

Next Generation Experimental Tracking for Machine Learning Operations

Category
model serving
Stars
391
Readiness
high risk (43/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
26/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, license. Risks: no push in 155 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

machine-learning mlops

NightMean/OlliteRT

Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source

Category
model serving
Stars
146
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

android anthropic-api gemma home-assistant kotlin-android litert

wladimiravila/esp32s3-distributed-ai

Distributed 56M-parameter LLM inference across 3 ESP32-S3 boards via ESP-NOW , Split-PLE + KV cache, fully offline.

Category
model serving
Stars
39
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

edge-ai edge-ai-engineering edge-ai-models embedded embedded-c embedded-systems

FareedKhan-dev/glm-5.2-in-c

GLM-5.2, a 744 billion parameter mixture of experts model, in a pure C inference engine: quantized to int4, experts streamed from disk, deployed and benchmarked. Generates in 16 GB of RAM.

Category
model serving
Stars
32
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +16 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

c consumer-hardware cpu-inference cuda deep-learning edge-ai

mit-han-lab/TinyChatEngine

TinyChatEngine: On-Device LLM Inference Library

Category
model serving
Stars
960
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 764 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

arm c cpp cuda-programming deep-learning edge-computing

unixsysdev/llama-turboquant

TurboQuant 3-bit KV-cache quantization for llama.cpp

Category
model serving
Stars
53
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal model serving project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

kv-cache llama-cpp llm-inference quantization turboquant