← OSS.Radar home
Powered by aegismemory.com · Aegis Memory repository
Evals And Testing AI repositories
OSS Radar projects in the evals and testing category.
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
- Category
- evals and testing
- Stars
- 32,715
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +467 stars in 7 days; 100+ commits in 30 days
Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals
Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.
Capped lower bounds: 30-day commits, lifetime contributors, response activity.
analytics autogen evaluation langchain large-language-models llama-index
Something wrong? Category · Trend · Risk
The missing DevTools for Claude Code — inspect session logs, tool calls, token usage, subagents, and context window in a visual UI. Free, open source.
- Category
- evals and testing
- Stars
- 3,804
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +24 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-agent ai-debugging ai-tools anthropic claude
Something wrong? Category · Trend · Risk
The LLM's practical guide: From the fundamentals to deploying advanced LLM and RAG apps to AWS using LLMOps best practices
- Category
- evals and testing
- Stars
- 5,271
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aws fine-tuning-llm genai llm llm-evaluation llmops
Something wrong? Category · Trend · Risk
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controllin
- Category
- evals and testing
- Stars
- 27,414
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +108 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentops agents ai ai-governance apache-spark evaluation
Something wrong? Category · Trend · Risk
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
- Category
- evals and testing
- Stars
- 21,198
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +192 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
evaluation hacktoberfest hacktoberfest2025 langchain llama-index llm
Something wrong? Category · Trend · Risk
AI Observability & Evaluation
- Category
- evals and testing
- Stars
- 10,939
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +104 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents ai-monitoring ai-observability aiengineering anthropic datasets
Something wrong? Category · Trend · Risk
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
- Category
- evals and testing
- Stars
- 6,045
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +22 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-monitoring analytics evaluation gpt langchain large-language-models
Something wrong? Category · Trend · Risk
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to
- Category
- evals and testing
- Stars
- 5,681
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +18 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent agent-evaluation agent-observability agentops ai coze
Something wrong? Category · Trend · Risk
🕹️ Open-source, developer-first LLMOps platform designed to streamline prompt design, version management, instant delivery, collaboration, troubleshooting, observability and more.
- Category
- evals and testing
- Stars
- 3,262
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai devtools gpt-3 gpt-4 hacktoberfest javascript
Something wrong? Category · Trend · Risk
Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Guardrails, Evaluations, Prompt Management, Vault, Playground. 🚀💻 Integrates with 50+ LLM Providers, Ve
- Category
- evals and testing
- Stars
- 2,674
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +13 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-observability amd-gpu clickhouse distributed-tracing genai gpu-monitoring
Something wrong? Category · Trend · Risk
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
- Category
- evals and testing
- Stars
- 1,055
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent agentic-ai agents grpo langchain langgraph
Something wrong? Category · Trend · Risk
Langtrace 🔍 is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vectorDB
- Category
- evals and testing
- Stars
- 1,225
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 263 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai datasets evaluations gpt langchain llm
Something wrong? Category · Trend · Risk
Multi-language agent runtime and library for execution scope management, lifecycle events, and middleware on tool and LLM calls.
- Category
- evals and testing
- Stars
- 113
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +26 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents ai atif atof claude-code codex
Something wrong? Category · Trend · Risk
an MLOps/LLMOps platform
- Category
- evals and testing
- Stars
- 237
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: documentation, license. Risks: no push in 595 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai cloud-native dataset datastore fine-tuning infra
Something wrong? Category · Trend · Risk
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
- Category
- evals and testing
- Stars
- 7,360
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artifical-intelligence datascience generative-ai good-first-issue good-first-issues help-wanted
Something wrong? Category · Trend · Risk
Ship AI Agents to Google Cloud in minutes, not months. Production-ready templates with built-in CI/CD, evaluation, and observability.
- Category
- evals and testing
- Stars
- 6,536
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents gcp gemini genai-agents generative-ai llmops
Something wrong? Category · Trend · Risk
AI observability platform for production LLM and agent systems.
- Category
- evals and testing
- Stars
- 4,416
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-observability ai ai-observability ai-tools evals fastapi
Something wrong? Category · Trend · Risk
Local-first AI token usage & cost tracker for 28 coding tools incl. Claude Code, Codex, Cursor, Gemini & Qoder—with native apps. Never reads prompts.
- Category
- evals and testing
- Stars
- 1,227
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +82 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-coding-tools ai-tools antigravity claude-code cli codex-cli
Something wrong? Category · Trend · Risk
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written in Python.
- Category
- evals and testing
- Stars
- 326
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai generative-ai linux-foundation llm-agent llm-inference llms
Something wrong? Category · Trend · Risk
The official evaluation suite and dynamic data release for MixEval.
- Category
- evals and testing
- Stars
- 254
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 635 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
benchmark benchmark-mixture benchmarking-framework benchmarking-suite evaluation evaluation-framework
Something wrong? Category · Trend · Risk
The platform for LLM evaluations and AI agent testing
- Category
- evals and testing
- Stars
- 3,479
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +38 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai analytics datasets dspy evaluation gpt
Something wrong? Category · Trend · Risk
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.
- Category
- evals and testing
- Stars
- 74
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai ai-agent-development ai-observability libpcap litellm llm-monitoring
Something wrong? Category · Trend · Risk
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses,
- Category
- evals and testing
- Stars
- 68
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
benchmarking chat-template cuda debugging llama-cpp llm-serving
Something wrong? Category · Trend · Risk
[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
- Category
- evals and testing
- Stars
- 1,188
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evaluation ai-safety confidence-estimation confidence-score hallucination hallucination-detection
Something wrong? Category · Trend · Risk
Collective Knowledge (CK), Collective Mind (CM/CMX) and MLPerf automations: community-driven projects to learn how to run AI, ML, and other emerging workloads more efficiently and cost-effectively acr
- Category
- evals and testing
- Stars
- 650
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation benchmarking best-practices ck cknowledge cm
Something wrong? Category · Trend · Risk
The LLM Evaluation Framework
- Category
- evals and testing
- Stars
- 17,470
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +164 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
evaluation-framework evaluation-metrics llm-evaluation llm-evaluation-framework llm-evaluation-metrics python
Something wrong? Category · Trend · Risk
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.
- Category
- evals and testing
- Stars
- 6,356
- Readiness
- ready (73/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai artificial-intelligence llm-agent llm-evaluation
Something wrong? Category · Trend · Risk
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
- Category
- evals and testing
- Stars
- 4,351
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agi audio-evaluation benchmark evaluation large-language-models llm-evaluation
Something wrong? Category · Trend · Risk
Evaluation and Tracking for LLM Experiments and AI Agents
- Category
- evals and testing
- Stars
- 3,491
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation agentops ai-agents ai-monitoring ai-observability evals
Something wrong? Category · Trend · Risk
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
- Category
- evals and testing
- Stars
- 3,150
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-observability agents ai ai-observability aiops analytics
Something wrong? Category · Trend · Risk
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
- Category
- evals and testing
- Stars
- 1,626
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +88 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents ai-evals ai-gateway ai-optimization ai-simulations evaluation-framework
Something wrong? Category · Trend · Risk
Prompty makes it easy to create, manage, debug, and evaluate LLM prompts for your AI applications. Prompty is an asset class and format for LLM prompts designed to enhance observability, understandab
- Category
- evals and testing
- Stars
- 1,246
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
generative-ai llm-evaluation llms promptengineering prompty
Something wrong? Category · Trend · Risk
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
- Category
- evals and testing
- Stars
- 1,021
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent financial-research llm-evaluation pgvector postgresql rabbitmq
Something wrong? Category · Trend · Risk
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
- Category
- evals and testing
- Stars
- 802
- Readiness
- ready (79/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +31 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation ai-agents awesome awesome-list benchmarks evals
Something wrong? Category · Trend · Risk
A test runner for agentskills.io-style AI agent skills
- Category
- evals and testing
- Stars
- 656
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evals agent-skills agentskills ai-agents cli jsonl
Something wrong? Category · Trend · Risk
Awesome papers involving LLMs in Social Science.
- Category
- evals and testing
- Stars
- 644
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alignment economics large-language-models llm-agent llm-evaluation llms
Something wrong? Category · Trend · Risk
Open-source benchmark for browser AI agents on daily tasks.
- Category
- evals and testing
- Stars
- 551
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +12 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation agentic-ai ai-agent-benchmark ai-agents benchmark browser-agent
Something wrong? Category · Trend · Risk
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
- Category
- evals and testing
- Stars
- 372
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents ci-cd evals llm-evaluation llm-observability llm-ops
Something wrong? Category · Trend · Risk
Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI frameworks and regulations. Join our Discord channel: https://discord.com/invite/d3k3E4
- Category
- evals and testing
- Stars
- 330
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-auditing ai-compliance ai-governance ai-governance-model ai-risk
Something wrong? Category · Trend · Risk
A list of LLMs Tools & Projects
- Category
- evals and testing
- Stars
- 322
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai chat-bot chatbots chatgpt data-science llm
Something wrong? Category · Trend · Risk
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
- Category
- evals and testing
- Stars
- 310
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio-codec benchmark evaluation llm-evaluation speech-recognition
Something wrong? Category · Trend · Risk
The definitive benchmark for AI agents on OpenClaw. 45 tasks across 4 tiers. Powered by MyClaw.ai
- Category
- evals and testing
- Stars
- 224
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-testing ai-agent ai-benchmark benchmark llm-evaluation myclaw
Something wrong? Category · Trend · Risk
A comprehensive guide to LLM evaluation methods designed to assist in identifying the most suitable evaluation techniques for various use cases, promote the adoption of best practices in LLM assessmen
- Category
- evals and testing
- Stars
- 198
- Readiness
- ready (78/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
evaluation generative-ai-benchmarking llm llm-benchmarking llm-evaluation
Something wrong? Category · Trend · Risk
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
- Category
- evals and testing
- Stars
- 162
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-as-judge benchmark computer-use-agent gui-agent hybrid-interface llm-evaluation
Something wrong? Category · Trend · Risk
Agent skill for long-form Chinese serial web-novel workflows with review gates, state guards, and quality signals.
- Category
- evals and testing
- Stars
- 145
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skill ai-writing chinese-webnovel codex llm-evaluation long-form-fiction
Something wrong? Category · Trend · Risk
开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation
- Category
- evals and testing
- Stars
- 116
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation ai-agent ai-evaluation ai-infra blind-test human-evaluation
Something wrong? Category · Trend · Risk
Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to local e
- Category
- evals and testing
- Stars
- 101
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation ai-evaluation evaluations infra llm-evaluation
Something wrong? Category · Trend · Risk
Evidence-driven multi-agent engineering harness: parallel agents, sealed evidence, independent juries, targeted rework, and traceable acceptance.
- Category
- evals and testing
- Stars
- 97
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation agent-governance agent-orchestration ai-agent-harness codex-cli coding-agents
Something wrong? Category · Trend · Risk
AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
- Category
- evals and testing
- Stars
- 81
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: fork interest, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents benchmark harbor llm-evaluation video-editing
Something wrong? Category · Trend · Risk
Comprehensive AI Model Evaluation Framework with advanced techniques including Temperature-Controlled Verdict Aggregation via Generalized Power Mean. Support for multiple LLM providers and 15+ evaluat
- Category
- evals and testing
- Stars
- 58
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evaluation ai-evaluation-framework ai-evaluation-metrics ai-evaluation-tools aieval llm-evaluation
Something wrong? Category · Trend · Risk
An Open Harness and Benchmark for AI in Cybersecurity Operations.
- Category
- evals and testing
- Stars
- 56
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai cybersecurity detection-engineering llm-agents llm-benchmark llm-evaluation
Something wrong? Category · Trend · Risk
discover LLMs punching above their weight
- Category
- evals and testing
- Stars
- 51
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
llm llm-evaluation local-llm
Something wrong? Category · Trend · Risk
Multi-Provider AI Gateway - No personal logs by design. Model autodiscovery, Failover groups, High availabilty, Android companion app, and more. "Because we have LiteLLM at home"
- Category
- evals and testing
- Stars
- 50
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-tools android-app discovery failover high-availability
Something wrong? Category · Trend · Risk
A design-of-experiments platform for evaluating compound AI systems - find which technique drives quality, by how much, and whether the difference is real.
- Category
- evals and testing
- Stars
- 45
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ab-testing compound-ai design-of-experiments evaluation factorial-design llm
Something wrong? Category · Trend · Risk
Causal Judge Evaluation: calibrate LLM-as-judge scores against oracle labels with valid uncertainty.
- Category
- evals and testing
- Stars
- 44
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evaluation calibration causal-inference evaluation-framework llm-as-judge llm-evaluation
Something wrong? Category · Trend · Risk
LLM Evaluation for Phoenix Apps
- Category
- evals and testing
- Stars
- 35
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai elixir live-view llm llm-evaluation phoenix
Something wrong? Category · Trend · Risk
End-to-end Langfuse workshop using a TypeScript Agent to teach the AI engineering loop: tracing, prompt management, monitoring, datasets, experiments, and evaluation.
- Category
- evals and testing
- Stars
- 29
- Readiness
- needs review (68/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-engineering developer-tools education langfuse llm-evaluation llm-observability
Something wrong? Category · Trend · Risk
TypeScript SDK for evaluating content quality using both traditional metrics and LLM-powered evaluation
- Category
- evals and testing
- Stars
- 27
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
eval evaluation evaluation-framework evaluation-metrics llm-evaluation
Something wrong? Category · Trend · Risk
pytest for LLM apps: record API calls once, replay them forever, and test meaning without flaky live runs.
- Category
- evals and testing
- Stars
- 25
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evals gena ghostrun llm llm-evaluation llm-testing
Something wrong? Category · Trend · Risk
Открытый бенчмарк LLM: какая нейросеть лучше пишет код 1С:Предприятие (BSL). Объективная оценка LLM по методике SMOP с реальным исполнением в 1С — Claude, GPT, Gemini, DeepSeek, YandexGPT, GigaChat.
- Category
- evals and testing
- Stars
- 24
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
1c 1c-enterprise agentic ai ai-code-generation benchmark
Something wrong? Category · Trend · Risk
The Regression Testing Framework for AI Agents. Replay · Evaluate · Assert · Catch Regressions — in CI. Like Jest for your AI layer.
- Category
- evals and testing
- Stars
- 23
- Readiness
- needs review (64/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-testing agentops ai-agent ai-agents ai-testing assertions
Something wrong? Category · Trend · Risk
A preservation-aware benchmark for natural-language SVG repair with deterministic binary rewards and complete model traces.
- Category
- evals and testing
- Stars
- 23
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evaluation benchmark llm-evaluation program-repair reproducible-research svg
Something wrong? Category · Trend · Risk
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
- Category
- evals and testing
- Stars
- 22
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artifact-validation azure-openai benchmark-automation dashboard gdpval github-actions
Something wrong? Category · Trend · Risk
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
- Category
- evals and testing
- Stars
- 19
- Readiness
- needs review (70/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation agent-evaluation-tools ai-agent-evaluation ai-agents ai-engineering ai-evals
Something wrong? Category · Trend · Risk
Real-world browser-agent benchmark: 210 tasks across 107 websites, multi-agent/multi-browser evaluation, reproducible leaderboard and result submissions.
- Category
- evals and testing
- Stars
- 19
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent agent-evaluation ai-agents benchmark browser-agent browser-automation
Something wrong? Category · Trend · Risk
Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule
- Category
- evals and testing
- Stars
- 18
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents ai ai-agents benchmark claude-code cli
Something wrong? Category · Trend · Risk
Expose what functional RTL benchmarks leave unanswered. Evidence profiles for AI-generated RTL; research collaborators and design partners welcome. Alpha research software, seeking validation
- Category
- evals and testing
- Stars
- 18
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-evaluation ai-hardware benchmark cdc clock-domain-crossing eda
Something wrong? Category · Trend · Risk
Production-grade LLM Evaluation & Benchmarking Framework — GPT-4, Claude, Gemini, Mistral. Accuracy, latency, cost, hallucination, reasoning metrics.
- Category
- evals and testing
- Stars
- 17
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accuracy ai benchmarking claude fastapi gemini
Something wrong? Category · Trend · Risk
Run a prompt against all, or some, of your models running on Ollama. Creates web pages with the output, performance statistics and model info. All in a single Bash shell script.
- Category
- evals and testing
- Stars
- 16
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-evaluation-tools attogram-project bash-script llm-eval llm-evaluation
Something wrong? Category · Trend · Risk
OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
- Category
- evals and testing
- Stars
- 16
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation ai benchmark bootstrap-ci claude claude-code
Something wrong? Category · Trend · Risk
Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
- Category
- evals and testing
- Stars
- 15
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ci ci-cd cicd evaluation evaluation-framework llm
Something wrong? Category · Trend · Risk
Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you swap models or edit prompts.
- Category
- evals and testing
- Stars
- 15
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-testing agents ai ai-agents ai-tools anthropic
Something wrong? Category · Trend · Risk
Reference harness for the Big Finance benchmark of workflow-grounded financial-research questions
- Category
- evals and testing
- Stars
- 15
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents ai-agents benchmark evaluation finance llm-evaluation
Something wrong? Category · Trend · Risk
No description
- Category
- evals and testing
- Stars
- 14
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility
Strongest signals: push recency, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
data-annotation llm llm-evaluation
Something wrong? Category · Trend · Risk
⚽🤖 11 frontier LLMs predicted the entire 2026 World Cup — frozen before kickoff. Live leaderboard: Brier score, bracket points & Polymarket ROI.
- Category
- evals and testing
- Stars
- 14
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai benchmark claude deepseek eval forecasting
Something wrong? Category · Trend · Risk
Build, enrich, and transform datasets using AI models with no code
- Category
- evals and testing
- Stars
- 1,637
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai llm-evaluation llms nocode oss synthetic-data
Something wrong? Category · Trend · Risk
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
- Category
- evals and testing
- Stars
- 654
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awsome-list awsome-lists benchmark bert chatglm chatgpt
Something wrong? Category · Trend · Risk
A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)
- Category
- evals and testing
- Stars
- 359
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 287 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
llm llm-agents llm-evaluation llmops llms-benchmarking
Something wrong? Category · Trend · Risk
AdaRubric: Adaptive Dynamic Rubric Evaluator for Agent Trajectories
- Category
- evals and testing
- Stars
- 348
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-evaluation llm-evaluation reward-model rlhf rubric
Something wrong? Category · Trend · Risk
Build, Improve Performance, and Productionize your AI Application
- Category
- evals and testing
- Stars
- 343
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 619 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai anthropic autogen docker full-stack javascript
Something wrong? Category · Trend · Risk
Python SDK for running evaluations on LLM generated responses
- Category
- evals and testing
- Stars
- 301
- Readiness
- high risk (41/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 427 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
evaluation evaluation-framework evaluation-metrics llm-eval llm-evaluation llm-evaluation-toolkit
Something wrong? Category · Trend · Risk
All-in-one Web Agent framework for post-training. Start building with a few clicks!
- Category
- evals and testing
- Stars
- 280
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 396 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent benchmark-framework llm-agent llm-evaluation
Something wrong? Category · Trend · Risk
Free, open-source AWS emulator. LocalStack alternative: 105 services, 7,391 operations, true 100% Smithy conformance (248,319/248,319 variants pass). No account, no auth token, no paid tier.
- Category
- evals and testing
- Stars
- 512
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-testing aws aws-bedrock aws-emulator aws-sdk aws-testing
Something wrong? Category · Trend · Risk
g4f-working is a daily-updated list of working no-auth AI providers and models from @xtekky/gpt4free. It helps developers, testers, and AI enthusiasts instantly find which models are currently online
- Category
- evals and testing
- Stars
- 130
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-access ai-directory ai-models ai-monitoring api auth-free
Something wrong? Category · Trend · Risk
A curated list of AI-powered testing tools, frameworks, and resources for QA engineers. From test generation to self-healing automation, MCP-based testing, LLM evaluation, and more.
- Category
- evals and testing
- Stars
- 52
- Readiness
- ready (78/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-test-automation ai-testing awesome awesome-list llm-testing mcp-testing
Something wrong? Category · Trend · Risk
🧪💥 Evaluation framework based on Vitest, the testing framework you familiar with, for agents, models, and more.
- Category
- evals and testing
- Stars
- 45
- Readiness
- needs review (65/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-evals evals evalscope evaluation evaluation-framework llm
Something wrong? Category · Trend · Risk
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
- Category
- evals and testing
- Stars
- 7
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-eval agent-evaluation ai-benchmarks ai-coding-agent-benchmark ai-engineering ai-evaluation
Something wrong? Category · Trend · Risk
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced anal
- Category
- evals and testing
- Stars
- 16,142
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 29/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 177 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai agentic-ai-development agentneo agents ai-agent-monitoring ai-application-debugging
Something wrong? Category · Trend · Risk
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
- Category
- evals and testing
- Stars
- 10
- Readiness
- needs review (68/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents ai-quality ai-testing amazon-bedrock eval-framework grounded-theory
Something wrong? Category · Trend · Risk
Unified CLI for running AI coding agents in isolated containers. Includes built-in local metrics collection, HTTP traffic tracking, and an analytics dashboard to track agent actions.
- Category
- evals and testing
- Stars
- 109
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai ai-agents ai-coding-assistant auggie-cli claude-code cli
Something wrong? Category · Trend · Risk
🦩 Tools for Go projects
- Category
- evals and testing
- Stars
- 4,497
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome awesome-list benchmarking code-generation code-visualization command-line-tool
Something wrong? Category · Trend · Risk
Open-source data management for multimodal AI. Query, trace, and govern data with a lineage-native lakehouse for files, tables, arrays, ontologies, and notes. With support for biological formats and r
- Category
- evals and testing
- Stars
- 280
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
comp-bio-ops context-engineering data-lakehouse data-lineage data-versioning eln
Something wrong? Category · Trend · Risk
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
- Category
- evals and testing
- Stars
- 23,573
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +63 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation data data-engineering data-ops data-science infrastructure
Something wrong? Category · Trend · Risk
First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.
- Category
- evals and testing
- Stars
- 1,420
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal evals and testing project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alerting bigdata data-catalog data-discovery data-engineering data-exploration
Something wrong? Category · Trend · Risk
📙 Awesome Data Catalogs and Observability Platforms.
- Category
- evals and testing
- Stars
- 1,056
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome awesome-list big-data data-catalog data-discovery data-engineering
Something wrong? Category · Trend · Risk