← OSS.Radar home

Powered by aegismemory.com · Aegis Memory repository

Evals And Testing AI repositories

OSS Radar projects in the evals and testing category.

langfuse/langfuse

🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23

Category
evals and testing
Stars
32,715
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
high
Maintainer health
healthy
Maintenance risk
0/100 · high confidence

Why: +467 stars in 7 days; 100+ commits in 30 days

Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals

Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.

Capped lower bounds: 30-day commits, lifetime contributors, response activity.

analytics autogen evaluation langchain large-language-models llama-index

matt1398/claude-devtools

The missing DevTools for Claude Code — inspect session logs, tool calls, token usage, subagents, and context window in a visual UI. Free, open source.

Category
evals and testing
Stars
3,804
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +24 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-agent ai-debugging ai-tools anthropic claude

PacktPublishing/LLM-Engineers-Handbook

The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices

Category
evals and testing
Stars
5,271
Readiness
needs review (62/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +11 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

aws fine-tuning-llm genai llm llm-evaluation llmops

mlflow/mlflow

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controllin

Category
evals and testing
Stars
27,414
Readiness
ready (97/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +108 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentops agents ai ai-governance apache-spark evaluation

comet-ml/opik

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

Category
evals and testing
Stars
21,198
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +192 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

evaluation hacktoberfest hacktoberfest2025 langchain llama-index llm

Arize-ai/phoenix

AI Observability & Evaluation

Category
evals and testing
Stars
10,939
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +104 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai-monitoring ai-observability aiengineering anthropic datasets

Helicone/helicone

🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

Category
evals and testing
Stars
6,045
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +22 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-monitoring analytics evaluation gpt langchain large-language-models

coze-dev/coze-loop

Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to

Category
evals and testing
Stars
5,681
Readiness
ready (94/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +18 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agent-evaluation agent-observability agentops ai coze

pezzolabs/pezzo

🕹️ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.

Category
evals and testing
Stars
3,262
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 30 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai devtools gpt-3 gpt-4 hacktoberfest javascript

openlit/openlit

Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, Ve

Category
evals and testing
Stars
2,674
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +13 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-observability amd-gpu clickhouse distributed-tracing genai gpu-monitoring

JudgmentLabs/judgeval

The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

Category
evals and testing
Stars
1,055
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agentic-ai agents grpo langchain langgraph

Scale3-Labs/langtrace

Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vectorDB

Category
evals and testing
Stars
1,225
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 263 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai datasets evaluations gpt langchain llm

NVIDIA/NeMo-Relay

Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.

Category
evals and testing
Stars
113
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +26 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai atif atof claude-code codex

star-whale/starwhale

an MLOps/LLMOps platform

Category
evals and testing
Stars
237
Readiness
needs review (48/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: limited evidence; inspect maintenance signals before adopting

Strongest signals: documentation, license. Risks: no push in 595 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai cloud-native dataset datastore fine-tuning infra

traceloop/openllmetry

Open-source observability for your GenAI or LLM application, based on OpenTelemetry

Category
evals and testing
Stars
7,360
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artifical-intelligence datascience generative-ai good-first-issue good-first-issues help-wanted

GoogleCloudPlatform/agent-starter-pack

Ship AI Agents to Google Cloud in minutes, not months. Production-ready templates with built-in CI/CD, evaluation, and observability.

Category
evals and testing
Stars
6,536
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents gcp gemini genai-agents generative-ai llmops

pydantic/logfire

AI observability platform for production LLM and agent systems.

Category
evals and testing
Stars
4,416
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-observability ai ai-observability ai-tools evals fastapi

xiufengsun/TokenTracker

Local-first AI token usage & cost tracker for 28 coding tools incl. Claude Code, Codex, Cursor, Gemini & Qoder—with native apps. Never reads prompts.

Category
evals and testing
Stars
1,227
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +82 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-coding-tools ai-tools antigravity claude-code cli codex-cli

monocle2ai/monocle

Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.

Category
evals and testing
Stars
326
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +6 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai generative-ai linux-foundation llm-agent llm-inference llms

JinjieNi/MixEval

The official evaluation suite and dynamic data release for MixEval.

Category
evals and testing
Stars
254
Readiness
needs review (45/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 635 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

benchmark benchmark-mixture benchmarking-framework benchmarking-suite evaluation evaluation-framework

langwatch/langwatch

The platform for LLM evaluations and AI agent testing

Category
evals and testing
Stars
3,479
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +38 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai analytics datasets dspy evaluation gpt

Netis/heron

Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.

Category
evals and testing
Stars
74
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +8 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai ai-agent-development ai-observability libpcap litellm llm-monitoring

Blackwellboy/model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses,

Category
evals and testing
Stars
68
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

benchmarking chat-template cuda debugging llama-cpp llm-serving

cvs-health/uqlm

[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"

Category
evals and testing
Stars
1,188
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evaluation ai-safety confidence-estimation confidence-score hallucination hallucination-detection

mlcommons/ck

Collective Knowledge (CK), Collective Mind (CM/CMX) and MLPerf automations: community-driven projects to learn how to run AI, ML, and other emerging workloads more efficiently and cost-effectively acr

Category
evals and testing
Stars
650
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automation benchmarking best-practices ck cknowledge cm

confident-ai/deepeval

The LLM Evaluation Framework

Category
evals and testing
Stars
17,470
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +164 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

evaluation-framework evaluation-metrics llm-evaluation llm-evaluation-framework llm-evaluation-metrics python

jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.

Category
evals and testing
Stars
6,356
Readiness
ready (73/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +19 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai artificial-intelligence llm-agent llm-evaluation

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Category
evals and testing
Stars
4,351
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agi audio-evaluation benchmark evaluation large-language-models llm-evaluation

truera/trulens

Evaluation and Tracking for LLM Experiments and AI Agents

Category
evals and testing
Stars
3,491
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +14 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation agentops ai-agents ai-monitoring ai-observability evals

lmnr-ai/lmnr

Laminar - open-source observability platform purpose-built for AI agents. YC S24.

Category
evals and testing
Stars
3,150
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +19 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-observability agents ai ai-observability aiops analytics

future-agi/future-agi

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

Category
evals and testing
Stars
1,626
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +88 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents ai-evals ai-gateway ai-optimization ai-simulations evaluation-framework

microsoft/prompty

Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandab

Category
evals and testing
Stars
1,246
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

generative-ai llm-evaluation llms promptengineering prompty

juanjuandog/FinSight-AI

AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.

Category
evals and testing
Stars
1,021
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agent financial-research llm-evaluation pgvector postgresql rabbitmq

benchflow-ai/awesome-evals

A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.

Category
evals and testing
Stars
802
Readiness
ready (79/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +31 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation ai-agents awesome awesome-list benchmarks evals

darkrishabh/agent-skills-eval

A test runner for agentskills.io-style AI agent skills

Category
evals and testing
Stars
656
Readiness
ready (88/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +9 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evals agent-skills agentskills ai-agents cli jsonl

ValueByte-AI/Awesome-LLM-in-Social-Science

Awesome papers involving LLMs in Social Science.

Category
evals and testing
Stars
644
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

alignment economics large-language-models llm-agent llm-evaluation llms

TIGER-AI-Lab/ClawBench

Open-source benchmark for browser AI agents on daily tasks.

Category
evals and testing
Stars
551
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +12 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation agentic-ai ai-agent-benchmark ai-agents benchmark browser-agent

Jwuthri/Tracely

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

Category
evals and testing
Stars
372
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents ci-cd evals llm-evaluation llm-observability llm-ops

verifywise-ai/verifywise

Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI frameworks and regulations. Join our Discord channel: https://discord.com/invite/d3k3E4

Category
evals and testing
Stars
330
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-auditing ai-compliance ai-governance ai-governance-model ai-risk

PetroIvaniuk/llms-tools

A list of LLMs Tools & Projects

Category
evals and testing
Stars
322
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai chat-bot chatbots chatgpt data-science llm

OpenBMB/UltraEval-Audio

Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation

Category
evals and testing
Stars
310
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

asr audio-codec benchmark evaluation llm-evaluation speech-recognition

LeoYeAI/myclaw-bench

The definitive benchmark for AI agents on OpenClaw. 45 tasks across 4 tiers. Powered by MyClaw.ai

Category
evals and testing
Stars
224
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-testing ai-agent ai-benchmark benchmark llm-evaluation myclaw

alopatenko/LLMEvaluation

A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the adoption of best practices in LLM assessmen

Category
evals and testing
Stars
198
Readiness
ready (78/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

evaluation generative-ai-benchmarking llm llm-benchmarking llm-evaluation

weavebench/WeaveBench

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Category
evals and testing
Stars
162
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-as-judge benchmark computer-use-agent gui-agent hybrid-interface llm-evaluation

zhougz520/novel-architect

Agent skill for long-form Chinese serial web-novel workflows with review gates, state guards, and quality signals.

Category
evals and testing
Stars
145
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-skill ai-writing chinese-webnovel codex llm-evaluation long-form-fiction

huangyiminghappy/ai-eval-platform

开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation

Category
evals and testing
Stars
116
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation ai-agent ai-evaluation ai-infra blind-test human-evaluation

evaleval/every_eval_ever

Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local e

Category
evals and testing
Stars
101
Readiness
ready (95/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation ai-evaluation evaluations infra llm-evaluation

Towow-ai/Flowness

Evidence-driven multi-agent engineering harness: parallel agents, sealed evidence, independent juries, targeted rework, and traceable acceptance.

Category
evals and testing
Stars
97
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation agent-governance agent-orchestration ai-agent-harness codex-cli coding-agents

PhiloLabs/agentic-vbench

AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?

Category
evals and testing
Stars
81
Readiness
ready (72/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: fork interest, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents benchmark harbor llm-evaluation video-editing

meshkovQA/Eval-ai-library

Comprehensive AI Model Evaluation Framework with advanced techniques including Temperature-Controlled Verdict Aggregation via Generalized Power Mean. Support for multiple LLM providers and 15+ evaluat

Category
evals and testing
Stars
58
Readiness
ready (72/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evaluation ai-evaluation-framework ai-evaluation-metrics ai-evaluation-tools aieval llm-evaluation

DeepTempo/socbench

An Open Harness and Benchmark for AI in Cybersecurity Operations.

Category
evals and testing
Stars
56
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai cybersecurity detection-engineering llm-agents llm-benchmark llm-evaluation

tolitius/cupel

discover LLMs punching above their weight

Category
evals and testing
Stars
51
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-evaluation local-llm

hugalafutro/model-hotel

Multi-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"

Category
evals and testing
Stars
50
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-tools android-app discovery failover high-availability

fabian-lu/Cafe

A design-of-experiments platform for evaluating compound AI systems - find which technique drives quality, by how much, and whether the difference is real.

Category
evals and testing
Stars
45
Readiness
ready (72/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ab-testing compound-ai design-of-experiments evaluation factorial-design llm

cimo-labs/cje

Causal Judge Evaluation: calibrate LLM-as-judge scores against oracle labels with valid uncertainty.

Category
evals and testing
Stars
44
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evaluation calibration causal-inference evaluation-framework llm-as-judge llm-evaluation

ccarvalho-eng/aludel

LLM Evaluation for Phoenix Apps

Category
evals and testing
Stars
35
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai elixir live-view llm llm-evaluation phoenix

langfuse/langfuse-workshop

End-to-end Langfuse workshop using a TypeScript Agent to teach the AI engineering loop: tracing, prompt management, monitoring, datasets, experiments, and evaluation.

Category
evals and testing
Stars
29
Readiness
needs review (68/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-engineering developer-tools education langfuse llm-evaluation llm-observability

loveholidays/eval-kit

TypeScript SDK for evaluating content quality using both traditional metrics and LLM-powered evaluation

Category
evals and testing
Stars
27
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

eval evaluation evaluation-framework evaluation-metrics llm-evaluation

parthmax2/ghostrun

pytest for LLM apps: record API calls once, replay them forever, and test meaning without flaky live runs.

Category
evals and testing
Stars
25
Readiness
ready (91/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evals gena ghostrun llm llm-evaluation llm-testing

genlab-1c/prism

Открытый бенчмарк LLM: какая нейросеть лучше пишет код 1С:Предприятие (BSL). Объективная оценка LLM по методике SMOP с реальным исполнением в 1С — Claude, GPT, Gemini, DeepSeek, YandexGPT, GigaChat.

Category
evals and testing
Stars
24
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

1c 1c-enterprise agentic ai ai-code-generation benchmark

1304674612/agentbench

The Regression Testing Framework for AI Agents. Replay · Evaluate · Assert · Catch Regressions — in CI. Like Jest for your AI layer.

Category
evals and testing
Stars
23
Readiness
needs review (64/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-testing agentops ai-agent ai-agents ai-testing assertions

yug-space/vector-edit-gym

A preservation-aware benchmark for natural-language SVG repair with deterministic binary rewards and complete model traces.

Category
evals and testing
Stars
23
Readiness
ready (80/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evaluation benchmark llm-evaluation program-repair reproducible-research svg

hyeonsangjeon/gdpval-realworks

Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.

Category
evals and testing
Stars
22
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

artifact-validation azure-openai benchmark-automation dashboard gdpval github-actions

ARTPARK-SAHAI-ORG/calibrate

Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

Category
evals and testing
Stars
19
Readiness
needs review (70/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation agent-evaluation-tools ai-agent-evaluation ai-agents ai-engineering ai-evals

lexmount/browseruse-agent-bench

Real-world browser-agent benchmark: 210 tasks across 107 websites, multi-agent/multi-browser evaluation, reproducible leaderboard and result submissions.

Category
evals and testing
Stars
19
Readiness
ready (87/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent agent-evaluation ai-agents benchmark browser-agent browser-automation

jfrog/agent-belt

Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule

Category
evals and testing
Stars
18
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai ai-agents benchmark claude-code cli

shsridhar-beep/svgap

Expose what functional RTL benchmarks leave unanswered. Evidence profiles for AI-generated RTL; research collaborators and design partners welcome. Alpha research software, seeking validation

Category
evals and testing
Stars
18
Readiness
ready (85/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-evaluation ai-hardware benchmark cdc clock-domain-crossing eda

vignesh2027/LLM-Evaluation-Framework

Production-grade LLM Evaluation & Benchmarking Framework — GPT-4, Claude, Gemini, Mistral. Accuracy, latency, cost, hallucination, reasoning metrics.

Category
evals and testing
Stars
17
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

accuracy ai benchmarking claude fastapi gemini

attogram/ollama-multirun

Run a prompt against all, or some, of your models running on Ollama. Creates web pages with the output, performance statistics and model info. All in a single Bash shell script.

Category
evals and testing
Stars
16
Readiness
ready (76/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai ai-evaluation-tools attogram-project bash-script llm-eval llm-evaluation

lizhiyao/oh-my-knowledge

OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.

Category
evals and testing
Stars
16
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation ai benchmark bootstrap-ci claude claude-code

Eval-core/evalcore

Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.

Category
evals and testing
Stars
15
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ci ci-cd cicd evaluation evaluation-framework llm

arthi-arumugam-git/whatbroke

Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you swap models or edit prompts.

Category
evals and testing
Stars
15
Readiness
ready (90/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-testing agents ai ai-agents ai-tools anthropic

Rogo-Technologies/big-finance-benchmark

Reference harness for the Big Finance benchmark of workflow-grounded financial-research questions

Category
evals and testing
Stars
15
Readiness
ready (93/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agents ai-agents benchmark evaluation finance llm-evaluation

AI4Bharat/Anudesh-Frontend

No description

Category
evals and testing
Stars
14
Readiness
needs review (62/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility

Strongest signals: push recency, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

data-annotation llm llm-evaluation

mverab/WorldCupBench

⚽🤖 11 frontier LLMs predicted the entire 2026 World Cup — frozen before kickoff. Live leaderboard: Brier score, bracket points & Polymarket ROI.

Category
evals and testing
Stars
14
Readiness
ready (82/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai benchmark claude deepseek eval forecasting

huggingface/aisheets

Build, enrich, and transform datasets using AI models with no code

Category
evals and testing
Stars
1,637
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai llm-evaluation llms nocode oss synthetic-data

onejune2018/Awesome-LLM-Eval

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

Category
evals and testing
Stars
654
Readiness
needs review (54/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awsome-list awsome-lists benchmark bert chatglm chatgpt

JonathanChavezTamales/llm-leaderboard

A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)

Category
evals and testing
Stars
359
Readiness
needs review (46/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 287 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

llm llm-agents llm-evaluation llmops llms-benchmarking

alphadl/AdaRubrics

AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories

Category
evals and testing
Stars
348
Readiness
needs review (55/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
watch
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-evaluation llm-evaluation reward-model rlhf rubric

palico-ai/palico-ai

Build, Improve Performance, and Productionize your AI Application

Category
evals and testing
Stars
343
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 619 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai anthropic autogen docker full-stack javascript

athina-ai/athina-evals

Python SDK for running evaluations on LLM generated responses

Category
evals and testing
Stars
301
Readiness
high risk (41/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation. Risks: no push in 427 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

evaluation evaluation-framework evaluation-metrics llm-eval llm-evaluation llm-evaluation-toolkit

iMeanAI/WebCanvas

All-in-one Web Agent framework for post-training. Start building with a few clicks!

Category
evals and testing
Stars
280
Readiness
needs review (50/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
30/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: no push in 396 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent benchmark-framework llm-agent llm-evaluation

faiscadev/fakecloud

Free, open-source AWS emulator. LocalStack alternative: 105 services, 7,391 operations, true 100% Smithy conformance (248,319/248,319 variants pass). No account, no auth token, no paid tier.

Category
evals and testing
Stars
512
Readiness
ready (81/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +4 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-testing aws aws-bedrock aws-emulator aws-sdk aws-testing

Free-AI-Things/g4f-working

g4f-working is a daily-updated list of working no-auth AI providers and models from @xtekky/gpt4free. It helps developers, testers, and AI enthusiasts instantly find which models are currently online

Category
evals and testing
Stars
130
Readiness
ready (84/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-access ai-directory ai-models ai-monitoring api auth-free

tugkanboz/awesome-ai-testing

A curated list of AI-powered testing tools, frameworks, and resources for QA engineers. From test generation to self-healing automation, MCP-based testing, LLM evaluation, and more.

Category
evals and testing
Stars
52
Readiness
ready (78/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +7 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-test-automation ai-testing awesome awesome-list llm-testing mcp-testing

vieval-dev/vieval

🧪💥 Evaluation framework based on Vitest, the testing framework you familiar with, for agents, models, and more.

Category
evals and testing
Stars
45
Readiness
needs review (65/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +3 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-evals evals evalscope evaluation evaluation-framework llm

linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Category
evals and testing
Stars
7
Readiness
ready (77/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agent-eval agent-evaluation ai-benchmarks ai-coding-agent-benchmark ai-engineering ai-evaluation

raga-ai-hub/RagaAI-Catalyst

Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced anal

Category
evals and testing
Stars
16,142
Readiness
needs review (61/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
29/100 · low confidence

Why: +1 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, fork interest, documentation. Risks: no push in 177 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai agentic-ai-development agentneo agents ai-agent-monitoring ai-application-debugging

aws-samples/sample-GEDD

Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.

Category
evals and testing
Stars
10
Readiness
needs review (68/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

ai-agents ai-quality ai-testing amazon-bedrock eval-framework grounded-theory

VibePod/vibepod-cli

Unified CLI for running AI coding agents in isolated containers. Includes built-in local metrics collection, HTTP traffic tracking, and an analytics dashboard to track agent actions.

Category
evals and testing
Stars
109
Readiness
ready (83/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +5 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

agentic-ai ai-agents ai-coding-assistant auggie-cli claude-code cli

nikolaydubina/go-recipes

🦩 Tools for Go projects

Category
evals and testing
Stars
4,497
Readiness
ready (86/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awesome awesome-list benchmarking code-generation code-visualization command-line-tool

laminlabs/lamindb

Open-source data management for multimodal AI. Query, trace, and govern data with a lineage-native lakehouse for files, tables, arrays, ontologies, and notes. With support for biological formats and r

Category
evals and testing
Stars
280
Readiness
ready (89/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

comp-bio-ops context-engineering data-lakehouse data-lineage data-versioning eln

PrefectHQ/prefect

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

Category
evals and testing
Stars
23,573
Readiness
ready (92/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: +63 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

automation data data-engineering data-ops data-science infrastructure

opendatadiscovery/odd-platform

First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

Category
evals and testing
Stars
1,420
Readiness
ready (72/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
healthy
Maintenance risk
0/100 · low confidence

Why: High-signal evals and testing project

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

alerting bigdata data-catalog data-discovery data-engineering data-exploration

opendatadiscovery/awesome-data-catalogs

📙 Awesome Data Catalogs and Observability Platforms.

Category
evals and testing
Stars
1,056
Readiness
needs review (53/100 heuristic points; not a probability)
Data confidence
low
Maintainer health
risky
Maintenance risk
0/100 · low confidence

Why: +2 stars in 7 days

Why it may be a gem: healthy maintenance and project fundamentals

Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.

awesome awesome-list big-data data-catalog data-discovery data-engineering