← OSS.Radar home
Powered by aegismemory.com · Aegis Memory repository
Multimodal Media AI repositories
OSS Radar projects in the multimodal media category.
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
- Category
- multimodal media
- Stars
- 102,101
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 8/100 · high confidence
Why: +1,309 stars in 7 days; 42 commits in 30 days
Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity; open issue backlog is stable or shrinking
Strongest signals: push recency, commit activity, contributor breadth. Risks: no pull-request review responses in 30 days. Missing inputs: None.
Capped lower bounds: response activity.
ai-video-generator content-creation ffmpeg instagram-reels llm python
Something wrong? Category · Trend · Risk
Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
- Category
- multimodal media
- Stars
- 69,693
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +406 stars in 7 days; 100+ commits in 30 days
Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals
Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.
Capped lower bounds: 30-day commits, lifetime contributors, response activity.
agent deepseek fine-tuning gemma gemma3 gpt-oss
Something wrong? Category · Trend · Risk
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
- Category
- multimodal media
- Stars
- 48,314
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +248 stars in 7 days; 100+ commits in 30 days
Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity
Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.
Capped lower bounds: 30-day commits, lifetime contributors, response activity.
agents ai api audio-generation decentralized distributed
Something wrong? Category · Trend · Risk
A generative speech model for daily dialogue.
- Category
- multimodal media
- Stars
- 39,748
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- watch
- Maintenance risk
- 28/100 · high confidence
Why: +31 stars in 7 days; 57 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: contributor breadth, issue load, documentation. Risks: no push in 119 days, no pull-request review responses in 30 days. Missing inputs: None.
agent chat chatgpt chattts chinese chinese-language
Something wrong? Category · Trend · Risk
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (
- Category
- multimodal media
- Stars
- 28,443
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- watch
- Maintenance risk
- 18/100 · high confidence
Why: +784 stars in 7 days; 11 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: contributor breadth, issue load, documentation. Risks: recent commit cadence is 0/20.0 of its monthly baseline. Missing inputs: None.
Capped lower bounds: response activity.
ai ai-meeting-assistant llm local-ai mac meeting-minutes
Something wrong? Category · Trend · Risk
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
- Category
- multimodal media
- Stars
- 12,941
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +17 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 276 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cogvideox image-to-video llm sora text-to-video video-generation
Something wrong? Category · Trend · Risk
AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.
- Category
- multimodal media
- Stars
- 23,464
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +592 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent claude-code codex html-deck image-generation ppt
Something wrong? Category · Trend · Risk
🧠 Leon is your open-source personal assistant.
- Category
- multimodal media
- Stars
- 17,419
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +23 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-agent ai-assistant artificial-intelligence assistant automation
Something wrong? Category · Trend · Risk
AI Agent 驱动的开源可自部署视频工作台:将小说与剧本转为角色、场景、道具资产、分镜、视频和剪映草稿,支持跨镜头一致性、多供应商与费用追踪 | Self-hosted AI video workspace for stories, storyboards and short-form video production
- Category
- multimodal media
- Stars
- 3,906
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +111 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-workflow ai-agent ai-animation ai-video-generator capcut claude-agent-sdk
Something wrong? Category · Trend · Risk
中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill
- Category
- multimodal media
- Stars
- 9,158
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +325 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent chinese codex-skill handdrawn illustration image-generation
Something wrong? Category · Trend · Risk
🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-file HTML → PNG. 小红书图文 + 公众号封面对
- Category
- multimodal media
- Stars
- 6,047
- Readiness
- needs review (59/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +327 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skill ai-agent anthropic claude-code claude-skill codex
Something wrong? Category · Trend · Risk
Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An
- Category
- multimodal media
- Stars
- 4,278
- Readiness
- needs review (64/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +48 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent apache-2 coding-agent css ffmpeg html
Something wrong? Category · Trend · Risk
中文手绘技术 PPT 整页图像生成 Skill | 21:9 封面 + 16:9 正文配图 | PNG 输出
- Category
- multimodal media
- Stars
- 1,310
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 17/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 104 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent chinese codex-skill handdrawn image-generation ppt
Something wrong? Category · Trend · Risk
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
- Category
- multimodal media
- Stars
- 1,556
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-generation image-to-video mcp mcp-server mcp-tools text-to-image
Something wrong? Category · Trend · Risk
Talk to your Mac, query your docs, no cloud required. On-device voice AI + RAG
- Category
- multimodal media
- Stars
- 1,540
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-assistant apple-silicon kitten-tts kokoro-tts lfm2 llama-cpp
Something wrong? Category · Trend · Risk
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
- Category
- multimodal media
- Stars
- 740
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio cross-lingual deep-learning fine-tuning multi-lingual python
Something wrong? Category · Trend · Risk
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
- Category
- multimodal media
- Stars
- 22,635
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- watch
- Maintenance risk
- 18/100 · high confidence
Why: +119 stars in 7 days; 43 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: contributor breadth, issue load, documentation. Risks: recent commit cadence is 0/2.0 of its monthly baseline. Missing inputs: release recency.
Capped lower bounds: response activity.
audio-generation cantonese chatbot chatgpt chinese cosyvoice
Something wrong? Category · Trend · Risk
The cleanest, responsive AI workspace. Run highly tuned image and video workflows flawlessly from your desktop or your phone. Built on ComfyUI.
- Category
- multimodal media
- Stars
- 237
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +44 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art comfyui image-generation lora mixstudio mobile-first
Something wrong? Category · Trend · Risk
AI filmmaking on a node canvas. Generate locally on your own GPU with the Inline Core engine (Z-Image, Krea 2, FLUX.2, MiniMax H3), or use hosted models with no GPU. Train your own LoRAs locally on th
- Category
- multimodal media
- Stars
- 198
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +16 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art ai-film ai-filmmaking ai-video controlnet creative-tools
Something wrong? Category · Trend · Risk
Demonstration for the Qwen-Image-Edit-2511 model with lazy-loaded LoRA adapters for advanced single- and multi-image editing. Supports 7+ specialized LoRAs including photo-to-anime, multi-angle camera
- Category
- multimodal media
- Stars
- 117
- Readiness
- ready (99/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusers flash-attention-3 huggingface-transformers image-editor image-generation image-to-image
Something wrong? Category · Trend · Risk
A microframework on top of PyTorch with first-class citizen APIs for foundation model adaptation
- Category
- multimodal media
- Stars
- 837
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 324 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
background-generation background-removal controlnet diffusion-models dinov2 image-generation
Something wrong? Category · Trend · Risk
In-context subject-driven image generation while preserving foreground fidelity
- Category
- multimodal media
- Stars
- 351
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 422 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models flux image-generation img2img lora subject-driven-generation
Something wrong? Category · Trend · Risk
Unrestricted Open-source alternative to AI video platforms — Free AI image & video generation studio with 500+ models (Flux, Midjourney, Kling, Sora, Veo). No content filters. Self-hosted, MIT license
- Category
- multimodal media
- Stars
- 25,808
- Readiness
- ready (96/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +529 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art-generator ai-image-generation ai-video-generation creative-tools fal-ai-alternative flux-1
Something wrong? Category · Trend · Risk
Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source AI tool that turns stories and scripts into animated short dramas. Features AI sc
- Category
- multimodal media
- Stars
- 13,536
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +387 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-content-creation ai-tool ai-video-generation automation content-generation
Something wrong? Category · Trend · Risk
Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.
- Category
- multimodal media
- Stars
- 3,998
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +51 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills agent-tools ai-agents ai-video claude-code claude-code-skills
Something wrong? Category · Trend · Risk
✨ Reverse-engineered Python API for Google Gemini web app
- Category
- multimodal media
- Stars
- 3,379
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai api async bard chatbot gemini
Something wrong? Category · Trend · Risk
A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可
- Category
- multimodal media
- Stars
- 3,337
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +518 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-filmmaking ai-video aigc aigc-pipeline content-creation
Something wrong? Category · Trend · Risk
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, Au
- Category
- multimodal media
- Stars
- 3,230
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +10 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ace-step ai audio-generation cosyvoice generative-ai generator
Something wrong? Category · Trend · Risk
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *
- Category
- multimodal media
- Stars
- 8,724
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 270 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auto-regressive-model autoregressive-models diffusion-models generative-ai generative-model gpt
Something wrong? Category · Trend · Risk
OpenMMLab Multimodal Advanced, Generative, and Intelligent Creation Toolbox. Unlock the magic 🪄: Generative-AI (AIGC), easy-to-use APIs, awsome model zoo, diffusion models, for text-to-image generatio
- Category
- multimodal media
- Stars
- 7,450
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 731 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc computer-vision deep-learning diffusion diffusion-models generative-adversarial-network
Something wrong? Category · Trend · Risk
Turn any face into a video game character, pixel art, claymation, 3D or toy
- Category
- multimodal media
- Stars
- 1,363
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 850 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai cog comfyui generative-ai replicate text-to-image
Something wrong? Category · Trend · Risk
[ICCV 2023] "TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition" (Official Implementation)
- Category
- multimodal media
- Stars
- 814
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 519 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-model generative-ai image-composition image-inversion stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
[ICML 2024] MagicPose(also known as MagicDance): Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
- Category
- multimodal media
- Stars
- 777
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 765 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
behavior-generation cartoon-animation diffusion-models generative-ai generative-model image-editing
Something wrong? Category · Trend · Risk
[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
- Category
- multimodal media
- Stars
- 706
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 402 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models generative-ai image-to-video transformer video-generation
Something wrong? Category · Trend · Risk
50+ open-source generative AI apps you can clone, deploy, and monetize — image generators, video tools, virtual try-ons, AI SaaS templates, and platform integrations. One-click Vercel deploy on every
- Category
- multimodal media
- Stars
- 2,852
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +55 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-apps ai-image-generator ai-saas ai-tools ai-video-generator awesome
Something wrong? Category · Trend · Risk
WebAI2API: 基于 Camoufox 的网页 AI 转 API 工具,支持 LMArena/Gemini等,多窗口并发与账号隔离。 | Web AI to OpenAI API via Camoufox. Supports LMArena/Gemini and more, multi-window concurrency & account isolation.
- Category
- multimodal media
- Stars
- 1,212
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-tools browser-automation generative-ai image-generation openai-api text-generation
Something wrong? Category · Trend · Risk
Official CLI for muapi.ai — generate images, videos & audio from the terminal. MCP server, 14 AI models, npm + pip installable.
- Category
- multimodal media
- Stars
- 1,037
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-api ai-tools api-client audio-generation cli
Something wrong? Category · Trend · Risk
Generate video from text using AI
- Category
- multimodal media
- Stars
- 795
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-api ai-tools ai-video ai-video-generator artificial-intelligence generative-ai
Something wrong? Category · Trend · Risk
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
- Category
- multimodal media
- Stars
- 1,685
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
computer-vision executorch image-embeddings llm-inference object-detection ocr
Something wrong? Category · Trend · Risk
🍌 World's largest Nano Banana Pro prompt library — 10,000+ curated prompts with preview images, 16 languages. Google Gemini AI image generation. Free & open source.
- Category
- multimodal media
- Stars
- 13,086
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +68 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image-generation ai-prompts awesome awesome-list gemini gemini-ai
Something wrong? Category · Trend · Risk
🚀 World's largest GPT Image 2 prompt library, updated daily — 2000+ curated prompts with preview images, 16 languages. OpenAI's next-gen image model with pixel-perfect text rendering, cross-image cons
- Category
- multimodal media
- Stars
- 9,175
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +171 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image-generation ai-prompts awesome awesome-list commercial-illustration duct-tape
Something wrong? Category · Trend · Risk
This repository contains a hand-curated resources for Prompt Engineering with a focus on Generative Pre-trained Transformer (GPT), ChatGPT, PaLM etc
- Category
- multimodal media
- Stars
- 6,234
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +22 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chatgpt chatgpt-api deep-learning few-shot-learning gpt gpt-3
Something wrong? Category · Trend · Risk
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
- Category
- multimodal media
- Stars
- 2,979
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +68 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-prompt asr dictation linux llm macos
Something wrong? Category · Trend · Risk
AI skill for OpenClaw & Claude Code — recommend from 10000+ Nano Banana Pro (Gemini) image prompts. Smart search by use case, content remix, sample images.
- Category
- multimodal media
- Stars
- 1,817
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +26 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-image claude-code-skill clawhub content-creation gemini
Something wrong? Category · Trend · Risk
🎬 2000+ curated Seedance 2.0 video generation prompts — cinematic, anime, UGC, ads, meme styles. Includes Seedance API guides, character consistency tips, and advanced video workflows.
- Category
- multimodal media
- Stars
- 1,799
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +64 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video awesome awesome-list prompt-engineering seedance seedance-2
Something wrong? Category · Trend · Risk
Awesome curated collection of images and prompts generated by GPT-4o and gpt-image-1. Explore AI generated visuals created with ChatGPT and Sora, showcasing OpenAI’s advanced image generation capabili
- Category
- multimodal media
- Stars
- 8,116
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 18/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 438 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art ai-image-examples anime-ai-art awesome-list cartoon-style curated-collection
Something wrong? Category · Trend · Risk
[CVPR 2026] PromptEnhancer is a prompt-rewriting tool, refining prompts into clearer, structured versions for better image generation.
- Category
- multimodal media
- Stars
- 3,748
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
hunyuan hunyuan-image image-editing image-to-image prompt prompt-engineering
Something wrong? Category · Trend · Risk
A large-scale text-to-image prompt gallery dataset based on Stable Diffusion
- Category
- multimodal media
- Stars
- 1,389
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 757 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art computer-vision image-generation prompt-engineering stable-diffusion
Something wrong? Category · Trend · Risk
1,400+ curated trending AI image prompts from X, ranked by engagement. Works with NanoBanana, GPT Image 2, Midjourney
- Category
- multimodal media
- Stars
- 706
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome-list gemini3proimage gpt-image image-generation midjourney nanobanana
Something wrong? Category · Trend · Risk
📚 GPT4o Prompts Dictionary | Curated Collection of AI Image Generation Prompts
- Category
- multimodal media
- Stars
- 583
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 18/100 · low confidence
Why: +3 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 450 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-image-creator ai-image-editing ai-image-generator ai-logo-maker aiimagegenerator
Something wrong? Category · Trend · Risk
One delightful Ruby framework for every major AI provider. Build AI agents, chatbots, RAG apps, and multimodal workflows in beautiful, expressive code.
- Category
- multimodal media
- Stars
- 4,286
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agents ai anthropic chatgpt claude deepseek
Something wrong? Category · Trend · Risk
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX bac
- Category
- multimodal media
- Stars
- 1,495
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +22 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
anthropic apple-silicon audio-processing claude-code computer-vision image-understanding
Something wrong? Category · Trend · Risk
Real-time speech-to-text WebSocket server with pluggable ASR backends, energy-based VAD, streaming partial results, and Prometheus observability.
- Category
- multimodal media
- Stars
- 90
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai asr qwen qwen-asr qwen3 qwen3-asr
Something wrong? Category · Trend · Risk
Gp.nvim (GPT prompt) Neovim AI plugin: ChatGPT sessions & Instructable text/code operations & Speech to text [OpenAI, Ollama, Anthropic, ..]
- Category
- multimodal media
- Stars
- 1,322
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 361 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
claude codeium copilot gemini gpt-4o gpt4o
Something wrong? Category · Trend · Risk
A real-time silent speech recognition tool.
- Category
- multimodal media
- Stars
- 749
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +5 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 278 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auto-avsr avsr llm ollama speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
- Category
- multimodal media
- Stars
- 8,719
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +85 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion dit linear-transformer nvfp4 pytorch reinforcement-learning
Something wrong? Category · Trend · Risk
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
- Category
- multimodal media
- Stars
- 7,694
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +32 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon audio-processing mlx multimodal speech-recognition speech-synthesis
Something wrong? Category · Trend · Risk
On-device Speech AI for Apple Silicon
- Category
- multimodal media
- Stars
- 6,311
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
inference ios macos pyannote qwen3-tts speaker-diarization
Something wrong? Category · Trend · Risk
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
- Category
- multimodal media
- Stars
- 1,295
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents architecture-diagrams arxiv awesome awesome-list computer-vision
Something wrong? Category · Trend · Risk
A PyTorch-based Speech Toolkit
- Category
- multimodal media
- Stars
- 11,743
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio audio-processing deep-learning huggingface language-model
Something wrong? Category · Trend · Risk
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
- Category
- multimodal media
- Stars
- 5,628
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 902 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanism deep-learning multi-modal text-to-image transformers
Something wrong? Category · Trend · Risk
Simple command line tool for text to image generation using OpenAI's CLIP and Siren (Implicit neural representation network). Technique was originally created by https://twitter.com/advadnoun
- Category
- multimodal media
- Stars
- 4,316
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1608 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning implicit-neural-representation multi-modality siren text-to-image
Something wrong? Category · Trend · Risk
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
- Category
- multimodal media
- Stars
- 2,836
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 332 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr attention-is-all-you-need attention-mechanism attention-model attention-network attention-seq2seq
Something wrong? Category · Trend · Risk
A playground to generate images from any text prompt using Stable Diffusion (past: using DALL-E Mini)
- Category
- multimodal media
- Stars
- 2,742
- Readiness
- needs review (60/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 795 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial artificial-intelligence dall-e dalle dalle-mini gan
Something wrong? Category · Trend · Risk
Text-to-Image generation. The repo for NeurIPS 2021 paper "CogView: Mastering Text-to-Image Generation via Transformers".
- Category
- multimodal media
- Stars
- 1,799
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1047 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
pretrained-models pytorch text-to-image transformers
Something wrong? Category · Trend · Risk
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
- Category
- multimodal media
- Stars
- 1,587
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 19/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 113 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auto-regressive-model autoregressive-models generative-model gpt gpt-2 image-generation
Something wrong? Category · Trend · Risk
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch
- Category
- multimodal media
- Stars
- 1,545
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 470 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanism audio-generation deep-learning non-autoregressive transformers
Something wrong? Category · Trend · Risk
Interface for OuteTTS models.
- Category
- multimodal media
- Stars
- 1,437
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 23/100 · low confidence
Why: +1 stars in 30 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 137 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
gguf llama text-to-speech transformers tts
Something wrong? Category · Trend · Risk
Generative Adversarial Transformers
- Category
- multimodal media
- Stars
- 1,346
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1515 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
attention compositionality gans generative-adversarial-networks image-generation scene-generation
Something wrong? Category · Trend · Risk
Self-hosted, OpenAI-compatible AI gateway for private RAG, natural-language data access, and tool-calling agents.
- Category
- multimodal media
- Stars
- 333
- Readiness
- ready (96/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-assistant ai-gateway ai-safety anthropic chatbot developer-tools
Something wrong? Category · Trend · Risk
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state mana
- Category
- multimodal media
- Stars
- 709
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-framework ai-voice ai-voice-agent audio-streaming golang open-source
Something wrong? Category · Trend · Risk
A framework for efficient model inference with omni-modality models
- Category
- multimodal media
- Stars
- 5,943
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +177 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-generation diffusion image-generation inference model-serving multimodal
Something wrong? Category · Trend · Risk
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
- Category
- multimodal media
- Stars
- 5,166
- Readiness
- ready (99/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +67 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-calling asterisk-ari conversational-ai inbound-calls local-llm no-code
Something wrong? Category · Trend · Risk
Custom TTS component for Home Assistant. Utilizes the OpenAI speech engine or any compatible endpoint to deliver high-quality speech. Optionally offers chime and audio normalization features.
- Category
- multimodal media
- Stars
- 204
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 16/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, fork interest, documentation. Risks: no push in 94 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai chime ha hacs home-assistant homeassistant
Something wrong? Category · Trend · Risk
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.
- Category
- multimodal media
- Stars
- 741
- Readiness
- ready (79/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-generation conversational-ai edge-ai lightweight llama multi-speaker
Something wrong? Category · Trend · Risk
On-device speech SDK for Android — ASR, TTS, VAD, and noise cancellation powered by ONNX Runtime with Qualcomm NNAPI acceleration
- Category
- multimodal media
- Stars
- 131
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android asr automotive edge-ai kotlin nnapi
Something wrong? Category · Trend · Risk
On-device VAD / streaming STT / TTS / diarization in C++17 (ONNX + LiteRT) with a voice-agent pipeline. Linux, Windows, Android.
- Category
- multimodal media
- Stars
- 64
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android cpp17 edge-ai kokoro litert nemotron
Something wrong? Category · Trend · Risk
edge-dit.cpp — a native C/C++ inference engine for Diffusion Transformers (DiT), designed for local and resource-constrained devices with automatic VRAM-aware precision, placement, and offloading.
- Category
- multimodal media
- Stars
- 25
- Readiness
- ready (99/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cpp diffusion-transformers dit edge-ai generative-ai ggml
Something wrong? Category · Trend · Risk
Deploy streaming ASR + TTS on RK3576/RK3588 — 120ms TTS, 52-language ASR, one Docker command
- Category
- multimodal media
- Stars
- 20
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr edge-ai npu rk3576 rk3588 rknn
Something wrong? Category · Trend · Risk
speech to text benchmark framework
- Category
- multimodal media
- Stars
- 697
- Readiness
- ready (73/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aws-transcribe cheetah deep-learning deep-neural-networks deepspeech edge-ai
Something wrong? Category · Trend · Risk
A practical lab for building, testing, and evaluating apps with Apple's Foundation Models framework.
- Category
- multimodal media
- Stars
- 1,165
- Readiness
- ready (79/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai apple-foundation-models apple-intelligence foundation-models foundation-models-framework generative-ai
Something wrong? Category · Trend · Risk
Muesli - local meeting transcription + dictation for macOS (Granola + WisprFlow alternative)
- Category
- multimodal media
- Stars
- 901
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +41 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon dictation macos meeting-assistant meeting-notes on-device-ai
Something wrong? Category · Trend · Risk
🎙️ VoxSherpa TTS Offline Neural Text-to-Speech Engine for Android ⚡ Sherpa-ONNX powered 🔊 Natural voice synthesis 📱 Fully offline processing 🚀 No cloud • No limits
- Category
- multimodal media
- Stars
- 192
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android android-ai android-app hindi-tts kokoro-82m kokoro-onnx
Something wrong? Category · Trend · Risk
PyTorch → CoreML conversion pipeline for Kokoro TTS. Unlocks fast on-device text-to-speech on Apple Neural Engine.
- Category
- multimodal media
- Stars
- 57
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-neural-engine apple-silicon on-device-ai text-to-speech
Something wrong? Category · Trend · Risk
On-device speech-to-text for macOS. Hold a hotkey → Whisper transcribes locally via CoreML → text auto-pastes into the focused app. 99 languages, free, open-source.
- Category
- multimodal media
- Stars
- 46
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon dictation macos on-device-ai productivity speech-to-text
Something wrong? Category · Trend · Risk
Hands-free on-device voice loop for macOS: Apple SFSpeechRecognizer + cloned-voice TTS, continuous listening, zero cloud. Originally built to drive claude-code-local.
- Category
- multimodal media
- Stars
- 45
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon claude-code hands-free local-ai macos offline-ai
Something wrong? Category · Trend · Risk
Open-source iOS voice dictation keyboard — fully offline, private, no subscription required for core features.
- Category
- multimodal media
- Stars
- 20
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
coreml dictation ios keyboard-extension offline on-device-ai
Something wrong? Category · Trend · Risk
Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacB
- Category
- multimodal media
- Stars
- 2,492
- Readiness
- needs review (65/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
computer-use-agents desktop-automation edge-computing gui-automation gui-grounding local-inference
Something wrong? Category · Trend · Risk
Generate images locally
- Category
- multimodal media
- Stars
- 525
- Readiness
- needs review (60/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
1-bit bonsai image-generation on-device-ai small-models ternary
Something wrong? Category · Trend · Risk
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
- Category
- multimodal media
- Stars
- 9,032
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +59 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio-analysis audio-event-detection cantonese cross-lingual emotion-detection
Something wrong? Category · Trend · Risk
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
- Category
- multimodal media
- Stars
- 4,049
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +215 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents amd comfyui docker llama-cpp llm
Something wrong? Category · Trend · Risk
Free offline AI video dubbing studio for Windows — voice cloning, translation, subtitles & on-screen-text localization. 100% local, one native .exe, zero Python.
- Category
- multimodal media
- Stars
- 87
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-dubbing ai-video asr dubbing llama-cpp
Something wrong? Category · Trend · Risk
Neve AI é uma plataforma de IA local privacy-first, desenvolvida para oferecer uma experiência de alta performance na execução de LLMs, reduzindo a dependência de grandes plataformas, assinaturas cara
- Category
- multimodal media
- Stars
- 49
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-workflow image-generation llama-cpp llm llms mtp
Something wrong? Category · Trend · Risk
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GP
- Category
- multimodal media
- Stars
- 38
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic diffusers embeddings gpu image-generation inference-server
Something wrong? Category · Trend · Risk
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video produc
- Category
- multimodal media
- Stars
- 45,916
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +1,554 stars in 7 days; 50 commits in 30 days
Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals
Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: release recency.
agent agentic-ai ai claude copilot cursor
Something wrong? Category · Trend · Risk
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models (LLM-grounded Diffusion: LMD, TMLR 2024)
- Category
- multimodal media
- Stars
- 484
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 697 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-generation llm stable-diffusion stable-diffusion-webui text-to-image
Something wrong? Category · Trend · Risk
An SDK/Python library for Automatic 1111 to run state-of-the-art diffusion models
- Category
- multimodal media
- Stars
- 412
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 793 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-art api automatic1111 deep-learning diffusers
Something wrong? Category · Trend · Risk
Official Agnes AI gateway and model catalog for OpenAI-compatible text, image, video, and agent workflows.
- Category
- multimodal media
- Stars
- 2,694
- Readiness
- needs review (69/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +504 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agnes-ai ai-api free-api multimodal-ai
Something wrong? Category · Trend · Risk
EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.
- Category
- multimodal media
- Stars
- 1,304
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +154 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aios multimodal-ai real-time robotics smart-devices voice-assistant
Something wrong? Category · Trend · Risk
Your Open Autonomous Android Agent — A production-ready, self-planning AI assistant powered by local/remote LLMs and accessibility-driven screen automation.
- Category
- multimodal media
- Stars
- 571
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +36 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility ai-agent android automation autonomous kotlin
Something wrong? Category · Trend · Risk
[ICML 2026] ByteDance's All-in-One Video Generation Model for Human-Object Interaction Video Generation
- Category
- multimodal media
- Stars
- 464
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- watch
- Maintenance risk
- 0/100 · high confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: maintenance is concentrated in one contributor. Missing inputs: release recency.
aigc computer-vision deep-learning diffusion-models dit icml
Something wrong? Category · Trend · Risk
MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.
- Category
- multimodal media
- Stars
- 409
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
llms long-video-understanding multimodal-ai real-time-ai streaming-video video-understanding
Something wrong? Category · Trend · Risk
Seedance 2.5 API guide, prompts, parameters, and examples for video generation
- Category
- multimodal media
- Stars
- 282
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video ai-video-generation bytedance camera-control cinematic-ai generative-ai
Something wrong? Category · Trend · Risk
Python wrapper for Black Forest Labs' FLUX 3 Dev API, available now — fast, low-cost FLUX 3 variant, plus the full FLUX 3 family (Text-to-Image, Image-to-Image, Text-to-Video, Image-to-Video).
- Category
- multimodal media
- Stars
- 176
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image ai-image-generation ai-video-generation black-forest-labs black-forest-labs-flux flux-3
Something wrong? Category · Trend · Risk
FLUX 3 video API and image API guide, prompts, parameters, and examples for Black Forest Labs' unified text-to-video, image-to-video, image, and audio generation model
- Category
- multimodal media
- Stars
- 142
- Readiness
- ready (78/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image-generation ai-video ai-video-generation api black-forest-labs flux
Something wrong? Category · Trend · Risk
Server-side video workflows for agents: ingest, understand, search, edit, stream.
- Category
- multimodal media
- Stars
- 115
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai amp claude claude-code codex multimodal-ai
Something wrong? Category · Trend · Risk
A modern multimodal knowledge graph with type-specific metadata across biomedical domains.
- Category
- multimodal media
- Stars
- 110
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
biomedical graph-ai heterogeneous-graphs knowledge-graph multimodal-ai multimodal-data
Something wrong? Category · Trend · Risk
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Stream
- Category
- multimodal media
- Stars
- 93
- Readiness
- ready (78/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents computer-vision lifelong-learning long-term-memory multimodal-ai openclaw
Something wrong? Category · Trend · Risk
🎬 Curated MiniMax H3 video generation prompts — cinematic, ads, anime, UGC, product videos, and more. Includes playable examples and creator attribution.
- Category
- multimodal media
- Stars
- 68
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video awesome-list generative-ai image-to-video minimax-h3 multimodal-ai
Something wrong? Category · Trend · Risk
基于多模态视觉感知与 LLM Agent 的 macOS 微信自动化框架 | Visual RPA for WeChat
- Category
- multimodal media
- Stars
- 65
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent ai-agent ai-automation apple-script automation chatbot
Something wrong? Category · Trend · Risk
LightMem-Ego: Your AI Memory for Everyday Life
- Category
- multimodal media
- Stars
- 64
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-memory artificial-intelligence lightmem lightmem-ego multimodal-ai
Something wrong? Category · Trend · Risk
Open-source, AI-enhanced CAT tool with multi-LLM support, translation memory, glossary management, 'Superlookup' concordance across TMs/glossaries/web resources, voice commands, Okapi sidecar for file
- Category
- multimodal media
- Stars
- 50
- Readiness
- ready (79/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ahk ai cafetran cat-tool claude context-aware-translation
Something wrong? Category · Trend · Risk
Evaluation tools and experiments for multimodal factual grounding, entity consistency, and image-text verification.
- Category
- multimodal media
- Stars
- 35
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
evaluation factual-grounding image-text multimodal-ai vision-language
Something wrong? Category · Trend · Risk
Private local AI Photographer Agent on AMD Radeon and ROCm
- Category
- multimodal media
- Stars
- 28
- Readiness
- ready (74/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent amd-rocm multimodal-ai photography qwen2-vl
Something wrong? Category · Trend · Risk
Multimodal AI SDK for IoT Devices
- Category
- multimodal media
- Stars
- 27
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent ai embedded iot multimodal-ai sdk
Something wrong? Category · Trend · Risk
Turn videos and courses into evidence-grounded Agent Skills for Claude Code and Codex.
- Category
- multimodal media
- Stars
- 21
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills claude-code codex multimodal-ai video youtube
Something wrong? Category · Trend · Risk
Open-source Android framework for low-latency, LLM-driven multimodal interaction on Pepper. Uses end-to-end speech-to-speech models and extensive Function Calling for agentic robot control (navigation
- Category
- multimodal media
- Stars
- 18
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android function-calling human-robot-interaction kotlin large-language-models multimodal-ai
Something wrong? Category · Trend · Risk
A unified, unmanaged multimodal runtime. Zero-allocation native AI inference for Java 22+.
- Category
- multimodal media
- Stars
- 15
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
foreign-function-and-memory-api ggml java-22 java-25 jvm llama-cpp
Something wrong? Category · Trend · Risk
Hermes Live Discord Agent Plugin — full-duplex Discord voice ↔ Google Gemini Multimodal Live API, with function calling, idle hangup, transcripts, and a 3-min oneshot installer.
- Category
- multimodal media
- Stars
- 14
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility
Strongest signals: push recency, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
discord-bot gemini-live hermes-agent hermes-plugin multimodal-ai plugin
Something wrong? Category · Trend · Risk
Demo of a Fiber Cut Response Agent using Azure Content Understanding to process multi-modal field documents (PDFs, photos, diagrams, audio) and reason with Foundry models for incident triage. From Mic
- Category
- multimodal media
- Stars
- 9
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents azure-ai-foundry azure-content-understanding build-2026 document-intelligence microsoft-build
Something wrong? Category · Trend · Risk
MedPMC: a high-fidelity medical image–text curation framework from biomedical literature for multimodal foundation models, with released pipelines, curated data, benchmarks, and pretrained models.
- Category
- multimodal media
- Stars
- 8
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
biomedical-literature dataset-curation foundation-models medical-ai medical-imaging multimodal-ai
Something wrong? Category · Trend · Risk
n8n community node for ByteDance Seedance 2.0 and Seedance 2 Mini — automate Text-to-Video, Image-to-Video, and video generation in n8n workflows.
- Category
- multimodal media
- Stars
- 7
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-automation ai-video bytedance generative-ai image-to-video low-code
Something wrong? Category · Trend · Risk
A framework for per-modality failure analysis and missing-data evaluation in multimodal clinical AI.
- Category
- multimodal media
- Stars
- 7
- Readiness
- needs review (67/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
clinical-ai failure-analysis healthcare-ai interpretability medical-imaging model-robustness
Something wrong? Category · Trend · Risk
Introducing Bimo 5 (Autonomous Multi-Modal AI Agent with real-time voice, vision processing, document analysis, image generation, and web search).
- Category
- multimodal media
- Stars
- 7
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-assistant ai-chat ai-workspace autonomous-agents flask
Something wrong? Category · Trend · Risk
Video OSINT agent: senses + OSINT reach for any agent.
- Category
- multimodal media
- Stars
- 7
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills multimodal-ai osint video
Something wrong? Category · Trend · Risk
Build and train a mini Kimi K3 from scratch — a pure-PyTorch playground from a custom 88M model to 2T-scale LLM profiles.
- Category
- multimodal media
- Stars
- 7
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning generative-ai kimi-k3 kv-cache llm llm-training
Something wrong? Category · Trend · Risk
Curated datasets, benchmarks, models, and tools for egocentric AI, embodied intelligence, VLA, world models, robotics, and wearable vision.
- Category
- multimodal media
- Stars
- 6
- Readiness
- needs review (63/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
augmented-reality awesome-list benchmarks computer-vision datasets ego4d
Something wrong? Category · Trend · Risk
P2P multi agent orchestration harness, universal automation system and IDE
- Category
- multimodal media
- Stars
- 6
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-swarms agentic-ai cpp crossplatform deep-research-agent go
Something wrong? Category · Trend · Risk
Multi-model AI chatbot built with React, Next.js, and Material UI, with image analysis and speech-to-text.
- Category
- multimodal media
- Stars
- 6
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-chatbot anthropic google-ai image-analysis material-ui multimodal-ai
Something wrong? Category · Trend · Risk
A Production-grade Multimodal AI System for Real-Time E-Commerce Review Analytics, Aspect Sentiment Analysis, Visual Defect Detection & Zero-Shot Product Reranking.
- Category
- multimodal media
- Stars
- 5
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
computer-vision deep-learning ecommerce machine-learning multimodal-ai natural-language-processing
Something wrong? Category · Trend · Risk
📄 Enable smart conversations with documents, images, and audio files using this advanced Retrieval-Augmented Generation system.
- Category
- multimodal media
- Stars
- 5
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
autogen automation cli document-processing dotnet gpt
Something wrong? Category · Trend · Risk
Production-ready multimodal RAG pipeline for PDFs. Extracts text, tables, and images down to the atomic element level via unstructured, chunks and AI-enriches them with a vision LLM, stores them in Ch
- Category
- multimodal media
- Stars
- 5
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chromadb multimodal-ai pdf-processing rag rag-pipeline streamlit-dashboard
Something wrong? Category · Trend · Risk
Extract and summarize social media content from platforms like Instagram, TikTok, X, and YouTube using Claude Code.
- Category
- multimodal media
- Stars
- 5
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai automation claude-code claude-skill content-analysis data-extraction
Something wrong? Category · Trend · Risk
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
- Category
- multimodal media
- Stars
- 14,377
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 18/100 · low confidence
Why: +97 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 108 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-avatar ai-avatars cloning cloning-tool digital-human multimodal-ai
Something wrong? Category · Trend · Risk
HEX is a whole-body vision-language-action framework for full-sized humanoid robots.
- Category
- multimodal media
- Stars
- 328
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cross-embodiment embodied-ai humanoid-robotics imitation-learning multimodal-ai vision-language-action-model
Something wrong? Category · Trend · Risk
个人相册语义资产生成器:扫描照片 → 视觉大模型标注 → 结构化数据 | Semantic asset generator for personal photo albums
- Category
- multimodal media
- Stars
- 240
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-captioning multimodal-ai openai-compatible photo-management
Something wrong? Category · Trend · Risk
GPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing
- Category
- multimodal media
- Stars
- 4,211
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +134 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills ai-image-prompts claude-code-skill cli codex-skill gpt-image
Something wrong? Category · Trend · Risk
LTX-Video Support for ComfyUI
- Category
- multimodal media
- Stars
- 4,038
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +24 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
comfyui diffusion-models dit image-to-video image-to-video-generation text-to-image
Something wrong? Category · Trend · Risk
FLUX, Stable Diffusion, SDXL, SD3, LoRA, Fine Tuning, DreamBooth, Training, Automatic1111, Forge WebUI, SwarmUI, DeepFake, TTS, Animation, Text To Video, Tutorials, Guides, Lectures, Courses, ComfyUI,
- Category
- multimodal media
- Stars
- 2,752
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art coding deepfake-generation dreambooth education flux-dev
Something wrong? Category · Trend · Risk
Fix AI pixel art and vector images right in your browser
- Category
- multimodal media
- Stars
- 889
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
browser-tools computer-vision pixel-art text-to-image vector-graphics
Something wrong? Category · Trend · Risk
Curated collection of reusable JSON prompt templates & style references for AI image generation. Updated daily.
- Category
- multimodal media
- Stars
- 541
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai design-resources generative-ai image-generation json midjourney
Something wrong? Category · Trend · Risk
AI Plugin is a powerful extension for the Payload CMS, integrating advanced AI capabilities to enhance content creation and management.
- Category
- multimodal media
- Stars
- 539
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-translate ai-writing ai-writing-tool content-generation gpt-image-1
Something wrong? Category · Trend · Risk
A Collection of Google Colab Notebooks for various projects
- Category
- multimodal media
- Stars
- 503
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
animate-x colab-notebooks comfyui-workflow flux-kontext framepack hidream
Something wrong? Category · Trend · Risk
Generate a video script, voice and a talking face completely with AI
- Category
- multimodal media
- Stars
- 477
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +5 stars in 7 days
Why it may be a gem: open issue backlog is stable or shrinking
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: release recency.
ai-video-generator artificial-intelligence faceless faceless-video image-to-video shorts
Something wrong? Category · Trend · Risk
Official implementation of AsymFlow, pi-Flow, GMFlow
- Category
- multimodal media
- Stars
- 459
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 8/100 · high confidence
Why: +1 stars in 7 days; 12 commits in 30 days
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: issue load, documentation, license. Risks: no pull-request review responses in 30 days. Missing inputs: None.
diffusion-models distillation flow-matching generative-ai generative-model image-generation
Something wrong? Category · Trend · Risk
[ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Learnin
- Category
- multimodal media
- Stars
- 411
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc bagel comfy comfyui diffusion image-editing
Something wrong? Category · Trend · Risk
attention map tools for huggingface/diffusers
- Category
- multimodal media
- Stars
- 409
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
attention-map cross-attention cross-attention-diffusers cross-attention-map diffusers huggingface
Something wrong? Category · Trend · Risk
Open-source AI video workbench. Bring any model or your local ComfyUI, and let Claude Code / Codex / Cursor direct it over MCP — storyboard, references, generation, editable first cut on a real timeli
- Category
- multimodal media
- Stars
- 399
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-video ai-video-generator bring-your-own-key comfyui creative-tools
Something wrong? Category · Trend · Risk
[CVPR 2024] "MACE: Mass Concept Erasure in Diffusion Models" (Official Implementation)
- Category
- multimodal media
- Stars
- 394
- Readiness
- needs review (68/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models generative-ai stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
[ICLR 2025] Official Implementation of Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
- Category
- multimodal media
- Stars
- 345
- Readiness
- needs review (69/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-image non-autoregressive-transformers text-to-image
Something wrong? Category · Trend · Risk
Create and customize your AI influencer open-source
- Category
- multimodal media
- Stars
- 279
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-api ai-influencer ai-tools artificial-intelligence creative-ai flux
Something wrong? Category · Trend · Risk
Code release for "i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models"
- Category
- multimodal media
- Stars
- 261
- Readiness
- needs review (68/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models jax pytorch text-to-image
Something wrong? Category · Trend · Risk
Modern AI image generator with multi-provider support (Gitee AI, HuggingFace, ModelScope), OpenAI-compatible API, token rotation, and one-click deployment to Cloudflare Pages.
- Category
- multimodal media
- Stars
- 216
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
biome cloudflare gitee-ai hono huggingface image-generation-ai
Something wrong? Category · Trend · Risk
Production-ready ComfyUI custom nodes for 1,400+ fal.ai models, auto-updated image, video, audio, LLM and VLM APIs with native media, caching and cost controls.
- Category
- multimodal media
- Stars
- 203
- Readiness
- ready (100/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-api audio-generation comfyui comfyui-custom-nodes fal fal-ai
Something wrong? Category · Trend · Risk
A curated list of recent style transfer methods with diffusion models
- Category
- multimodal media
- Stars
- 197
- Readiness
- needs review (67/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc diffusion-models style-transfer text-to-image text-to-video
Something wrong? Category · Trend · Risk
A front-end UI for ComfyUI made for beginner level users.
- Category
- multimodal media
- Stars
- 168
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art comfyui desktop-app flux generative-ai image-generation
Something wrong? Category · Trend · Risk
Local-first image prompt library for generating images, saving prompts, tags, and collections.
- Category
- multimodal media
- Stars
- 127
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image all-in-one allinone chatgpt fastapi gpt-image-2
Something wrong? Category · Trend · Risk
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.
- Category
- multimodal media
- Stars
- 93
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-skills ark cli doubao llm
Something wrong? Category · Trend · Risk
FastAPI wrapper for Meta AI with chat, image generation & video generation. Easy deployment with cookie-based auth. 🚀
- Category
- multimodal media
- Stars
- 90
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai api api-wrapper chat-api fastapi free
Something wrong? Category · Trend · Risk
Best GPT Image 2 OpenAi Prompts & Tools Guide 2026
- Category
- multimodal media
- Stars
- 88
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image chatgpt gpt-image-2 gpt-image-2-api gpt-image-2-prompts image-editing
Something wrong? Category · Trend · Risk
The most complete, up-to-date comparison of AI image generation models — which model, via which API, at what price.
- Category
- multimodal media
- Stars
- 78
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image ai-image-generator awesome awesome-list flux generative-ai
Something wrong? Category · Trend · Risk
A curated list of awesome stable diffusion resources 🌟
- Category
- multimodal media
- Stars
- 77
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aiart awesome dall-e diffusion diffusion-models dockerfile
Something wrong? Category · Trend · Risk
No description
- Category
- multimodal media
- Stars
- 74
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
api awsome-list gpt-image-2 gpt-image-2-prompts image-generation image-prompt
Something wrong? Category · Trend · Risk
OpenCode plugin: image generation via your ChatGPT subscription or OpenAI API
- Category
- multimodal media
- Stars
- 68
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents chatgpt codex gpt-image-2 image-generation openai
Something wrong? Category · Trend · Risk
ComfyUI custom nodes for 100+ AI models — Seedance, Kling, Veo3, Flux, HiDream, GPT-image, Imagen4 and more via muapi.ai
- Category
- multimodal media
- Stars
- 67
- Readiness
- ready (98/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image-editing ai-video-generation comfyui comfyui-custom-nodes comfyui-nodes flux
Something wrong? Category · Trend · Risk
Demonstrates Voice Recognition, Text to Speech, Language Translation, OAuth2, Image Generation, Face Detection and Voice Chatbot.
- Category
- multimodal media
- Stars
- 66
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai artificial-intelligence claude-3-haiku claude-3-opus claude-3-sonnet computer-vision
Something wrong? Category · Trend · Risk
⚡ CLI toolkit for Google Flow — Nano Banana Pro images, Omni Flash videos, MCP v2 & OpenAI API.
- Category
- multimodal media
- Stars
- 63
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation cli google-flow image-generation mcp nano-banana
Something wrong? Category · Trend · Risk
An interactive web demo for AI avatar creation — generate characters from a text prompt and inspect them in a 360° viewer. Powered by Gemini. 🍌🤖
- Category
- multimodal media
- Stars
- 52
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
360-viewer ai-art avatar-generator character-generation gemini generative-ai
Something wrong? Category · Trend · Risk
Tiny local text-to-24x24 pixel art model, trained on roughly 200K samples in 30 minutes on an RTX 5090.
- Category
- multimodal media
- Stars
- 52
- Readiness
- needs review (67/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +21 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
autoregressive-model fastapi generative-ai local-ai pixel-art pytorch
Something wrong? Category · Trend · Risk
Fully-local Discord bot with chat, RAG memory, autonomous Discord posts, Bluesky integration, and image generation. Runs against LM Studio + Stable Diffusion OR Flux, with a FastAPI web control panel
- Category
- multimodal media
- Stars
- 37
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
bot chatbot chatgpt-api discord discord-bot flux
Something wrong? Category · Trend · Risk
PortableLM — Zero-dependency, air-gapped local AI that runs entirely from a USB drive. Chat, generate images, and synthesize speech offline on Windows, macOS, or Linux. No installs, no internet, just
- Category
- multimodal media
- Stars
- 31
- Readiness
- ready (100/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-agents ai-model ai-tools image-generation llm
Something wrong? Category · Trend · Risk
Open-source Nano Banana image generator — production-ready Next.js SaaS for text-to-image and multi-image reference editing. Stripe billing, credits, NextAuth, and Prisma out of the box.
- Category
- multimodal media
- Stars
- 31
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art-generator ai-image-generator ai-photo-editor ai-saas gemini-2-5-flash gemini-image
Something wrong? Category · Trend · Risk
Text to Video API generation documentation
- Category
- multimodal media
- Stars
- 29
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video-generator image-to-video image-to-video-generation sora-ai sora-video-ai stable-diffusion
Something wrong? Category · Trend · Risk
Typescript - Lightweights WhatsApp bot 🤖 made to response only self message with Baileys ✨
- Category
- multimodal media
- Stars
- 28
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
baileys baileys-md bot converter downloader ffmpeg
Something wrong? Category · Trend · Risk
Private AI Image Generation Website Source Code powered by Gemini API | Ready-to-use | Text-to-Image/Image-to-Image/Multi-turn Chat | User System + Credits Billing + Redemption Codes | Built with Flas
- Category
- multimodal media
- Stars
- 28
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-image-generation flask gemini gemini-api generative-ai google-gemini
Something wrong? Category · Trend · Risk
Agent skill for OpenAI-compatible image generation, editing, and batch workflows.
- Category
- multimodal media
- Stars
- 25
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-tools batch-generation codex-skill image-generation openai-compatible python
Something wrong? Category · Trend · Risk
A full prompt-writing studio in a single ComfyUI node - AI editing, style targeting, prompt browsing, and wireless prompt injection, powered by your local Ollama models. No cloud, no keys.
- Category
- multimodal media
- Stars
- 24
- Readiness
- ready (74/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-prompt comfyui comfyui-custom-nodes comfyui-manager comfyui-nodes custom-nodes
Something wrong? Category · Trend · Risk
Kling AI Master
- Category
- multimodal media
- Stars
- 22
- Readiness
- needs review (70/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
kling kling-4 kling-ai text-to-image
Something wrong? Category · Trend · Risk
Use Atlas Cloud's 300+ AI models inside ComfyUI — drop-in nodes for Sora 2, Veo 3.1, Kling 3, Seedance 2, Nano Banana Pro, GPT Image 2, Flux 2 & more.
- Category
- multimodal media
- Stars
- 20
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai atlascloud comfyui comfyui-nodes generative-ai image-generation
Something wrong? Category · Trend · Risk
Experimental, entirely AI-coded ComfyUI nodes for MiniMax H3 T2I, I2I, reference editing, arbitrary frames, and optimized still selection.
- Category
- multimodal media
- Stars
- 19
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-assisted comfyui experimental generative-ai image-editing image-to-image
Something wrong? Category · Trend · Risk
小优AIGC(抖音小优与AIGC的奇妙冒险)的 AI 摄影剧组 skill —— 电影摄影级生图提示词生成器,可装入 Claude Code / Codex / Hermes
- Category
- multimodal media
- Stars
- 19
- Readiness
- needs review (70/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills aigc claude-code prompt-engineering text-to-image
Something wrong? Category · Trend · Risk
Local AI harness/workstation: local AI image editor, Ollama image generation routing, SDXL inpainting UI, CivitAI model imports, 8GB VRAM Stable Diffusion.
- Category
- multimodal media
- Stars
- 18
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
8gb-vram ai-assistant ai-harness ai-workstation chatgpt-alternative civitai
Something wrong? Category · Trend · Risk
A character forge for ComfyUI: dial in or randomize every detail - face, hair, body, wardrobe, mood - and get a coherent, reproducible person, canonically-described cosplayer, or creature as ready-to-
- Category
- multimodal media
- Stars
- 15
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
character-creator character-generator comfyui comfyui-custom-nodes prompt-generator text-to-image
Something wrong? Category · Trend · Risk
AI image generation CLI. One command, one image, three seconds. Gemini-powered.
- Category
- multimodal media
- Stars
- 15
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent-tools ai-art ai-image-generator cli command-line developer-tools
Something wrong? Category · Trend · Risk
Reproducible sketch-to-render studies with ControlNet: seed scouting, variation, export, replay, and a local Studio.
- Category
- multimodal media
- Stars
- 15
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
canny-edge-detection computer-vision controlnet diffusers generative-ai gradio
Something wrong? Category · Trend · Risk
Universal Prompt Generator — prompt-engineering system for text-to-image models. Spectrum of 1–5 calibrated prompts per theme, with safety/drift/cliche enforcement built into the pipeline.
- Category
- multimodal media
- Stars
- 14
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-prompts chatgpt claude flux gemini generative-ai
Something wrong? Category · Trend · Risk
A toy PyTorch implementation of FLUX diffusion transformers
- Category
- multimodal media
- Stars
- 14
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusion-transformer educational flux flux-architecture flux-kontext
Something wrong? Category · Trend · Risk
Budget-aware AI agent content studio for creating videos, images, carousels, voiceovers, music, captions, and content calendars with consistent brand assets and cost control.
- Category
- multimodal media
- Stars
- 13
- Readiness
- ready (96/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-connector ai ai-agents ai-video-toolkit claude-code content-creation-tools
Something wrong? Category · Trend · Risk
Awesome GPT Image 2 Prompts
- Category
- multimodal media
- Stars
- 12
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai awesome awesome-list generative-ai gpt-image-2 gpt-image-2-prompts
Something wrong? Category · Trend · Risk
AI image generation via 5 Chinese platforms — DashScope, Volcano Ark, Hunyuan, Zhipu, StepFun — 30 models, Chinese text excellence
- Category
- multimodal media
- Stars
- 12
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills ai-art alibaba-cloud claude-code claude-code-skill claude-skills
Something wrong? Category · Trend · Risk
轻量 GPT-Image-2 生图工作台,特别支持批量调用 GPT-Image-2:支持 AI 拆分提示词、队列、重试和图生图 / Lightweight GPT-Image-2 studio with batch calls, AI prompt splitting, queueing, retries, and image-to-image.
- Category
- multimodal media
- Stars
- 12
- Readiness
- needs review (71/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
desktop-app gpt-image-2 image-to-image local-first openai-compatible tauri
Something wrong? Category · Trend · Risk
475 tested Krea 2 Turbo prompts. One file, drop it in ComfyUI.
- Category
- multimodal media
- Stars
- 12
- Readiness
- ready (99/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art ai-prompts awesome awesome-list generative-ai image-generation
Something wrong? Category · Trend · Risk
Curated GPT-Image-2 prompts for the OpenAI API — portraits, posters, UI mockups, game screenshots, character sheets, and more. Ready-to-use prompts for gpt-image-2.
- Category
- multimodal media
- Stars
- 12
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art ai-generated-art ai-image api awesome-list chatgpt
Something wrong? Category · Trend · Risk
An open-source generator for AI image and text prompts that automatically builds richer, more detailed prompts than most people write by hand. Compose from a large library of scenes, subjects, and sty
- Category
- multimodal media
- Stars
- 11
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-art comfyui dynamic-prompts generative-ai image-generation
Something wrong? Category · Trend · Risk
Mcp server code to setup muapi to work with clients like Claude, Cursor etc.
- Category
- multimodal media
- Stars
- 10
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-api ai-tools anthropic api audio-generation
Something wrong? Category · Trend · Risk
Model-to-NPU pipelines for Qualcomm Snapdragon: QNN/ONNX/Android runtimes for on-device image and video generation.
- Category
- multimodal media
- Stars
- 10
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android diffusion-models hexagon-npu on-device-ai onnx qairt
Something wrong? Category · Trend · Risk
AI companions including generative AI such as chatbots, image generation, text generation, and audio generation.
- Category
- multimodal media
- Stars
- 9
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-translate chatbot generative-ai image-generation llm
Something wrong? Category · Trend · Risk
A .NET library for local image captioning using ONNX Runtime with automatic HuggingFace model download
- Category
- multimodal media
- Stars
- 9
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai csharp cuda directml dotnet flux
Something wrong? Category · Trend · Risk
Text-to-image web app: Stable Diffusion (Diffusers/PyTorch) behind a FastAPI backend that also serves a React frontend, with live streaming generation progress. Runs locally on GPU, CPU, or Apple MPS.
- Category
- multimodal media
- Stars
- 8
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusers fastapi pytorch react stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Krea 2 Turbo (12.9B text-to-image) ported to Apple MLX — local web UI + CLI, validated faithful to PyTorch
- Category
- multimodal media
- Stars
- 8
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon diffusion gradio image-generation krea mlx
Something wrong? Category · Trend · Risk
Run Krea 2 (Krea AI's 12B DiT) natively in Forge Neo — the first open-source Krea 2 extension. One-click model download, presets, fp8, full + piecewise loading. By Stable Yogi.
- Category
- multimodal media
- Stars
- 7
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art forge krea krea2 sd-webui-extension sd-webui-forge
Something wrong? Category · Trend · Risk
Implementation of DALL-E 2, OpenAI's updated text-to-image synthesis neural network, in Pytorch
- Category
- multimodal media
- Stars
- 11,307
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 818 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning text-to-image
Something wrong? Category · Trend · Risk
Implementation of Imagen, Google's Text-to-Image Neural Network, in Pytorch
- Category
- multimodal media
- Stars
- 8,418
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 669 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning imagination-machine text-to-image text-to-video
Something wrong? Category · Trend · Risk
Implementation of Dreambooth (https://arxiv.org/abs/2208.12242) with Stable Diffusion
- Category
- multimodal media
- Stars
- 7,737
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1338 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
pytorch pytorch-lightning stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Red Ink - A one-stop Xiaohongshu image-and-text generator based on the 🍌Nano Banana Pro🍌, "One Sentence, One Image: Generate Xiaohongshu Text and Images."
- Category
- multimodal media
- Stars
- 5,443
- Readiness
- needs review (64/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai aigc content-generator docker flask gemini
Something wrong? Category · Trend · Risk
min(DALL·E) is a fast, minimal port of DALL·E Mini to PyTorch
- Category
- multimodal media
- Stars
- 3,497
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 466 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning pytorch text-to-image
Something wrong? Category · Trend · Risk
Diffusion model papers, survey, and taxonomy
- Category
- multimodal media
- Stars
- 3,366
- Readiness
- high risk (41/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 314 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models stable-diffusion survey text-to-3d text-to-image text-to-video
Something wrong? Category · Trend · Risk
Kandinsky 2 — multilingual text2image latent diffusion model
- Category
- multimodal media
- Stars
- 2,815
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 828 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion image-generation image2image inpainting ipython-notebook kandinsky
Something wrong? Category · Trend · Risk
🔥 [ICCV 2025 Highlight] InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
- Category
- multimodal media
- Stars
- 2,685
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 350 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusers diffusion diffusion-transformer dit face flux
Something wrong? Category · Trend · Risk
Just playing with getting VQGAN+CLIP running locally, rather than having to use colab.
- Category
- multimodal media
- Stars
- 2,649
- Readiness
- high risk (42/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 1405 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
text-to-image text2image
Something wrong? Category · Trend · Risk
A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN. Technique was originally created by https://twitter.com/advadnoun
- Category
- multimodal media
- Stars
- 2,572
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1643 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning generative-adversarial-networks multimodality text-to-image
Something wrong? Category · Trend · Risk
(ෆ`꒳´ෆ) A Survey on Text-to-Image Generation/Synthesis.
- Category
- multimodal media
- Stars
- 2,441
- Readiness
- needs review (69/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awseome-list generative-adversarial-network image-generation image-manipulation image-synthesis multimodal
Something wrong? Category · Trend · Risk
AI magics meet Infinite draw board.
- Category
- multimodal media
- Stars
- 1,937
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 820 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-image inpainting latent-diffusion outpainting pypi python
Something wrong? Category · Trend · Risk
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
- Category
- multimodal media
- Stars
- 1,843
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 552 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-editting large-language-models multimodal-large-language-models text-to-image
Something wrong? Category · Trend · Risk
[ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"
- Category
- multimodal media
- Stars
- 1,740
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 598 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion diffusion-models eccv eccv2024 image-inpainting text-to-image
Something wrong? Category · Trend · Risk
Official Pytorch Implementation for "TokenFlow: Consistent Diffusion Features for Consistent Video Editing" presenting "TokenFlow" (ICLR 2024)
- Category
- multimodal media
- Stars
- 1,707
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 550 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
iclr2024 stable-diffusion text-to-image text-to-video tokenflow video-editing
Something wrong? Category · Trend · Risk
Generate images from texts. In Russian
- Category
- multimodal media
- Stars
- 1,645
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1305 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
dalle image-generation openai python pytorch russian
Something wrong? Category · Trend · Risk
[ICCV 2025] 🔥🔥 UNO: A Universal Customization Method for Both Single and Multi-Subject Conditioning
- Category
- multimodal media
- Stars
- 1,361
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 329 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion diffusion-transformer flux image-generation in-context-learning subject-driven-generation
Something wrong? Category · Trend · Risk
Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows
- Category
- multimodal media
- Stars
- 1,312
- Readiness
- needs review (63/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-art art asset-generator chatbot deep-learning
Something wrong? Category · Trend · Risk
A collection of resources on controllable generation with text-to-image diffusion models.
- Category
- multimodal media
- Stars
- 1,110
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 24/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 584 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome awesome-list controllable-generation diffusion-models multi-concept personalization
Something wrong? Category · Trend · Risk
CogView4, CogView3-Plus and CogView3(ECCV 2024)
- Category
- multimodal media
- Stars
- 1,100
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 496 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
eccv2024 high-resolution image-generation text-to-image
Something wrong? Category · Trend · Risk
Text2Room generates textured 3D meshes from a given text prompt using 2D text-to-image models (ICCV2023).
- Category
- multimodal media
- Stars
- 1,087
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 996 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
3d-generation diffusion-models mesh-generation text-to-image
Something wrong? Category · Trend · Risk
Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
- Category
- multimodal media
- Stars
- 1,065
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 1051 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models generative-model icml image-generation multidiffusion stable-diffusion
Something wrong? Category · Trend · Risk
Reverse-engineered the official API for Jimeng/Dreamina’s text-to-image and image-to-image features. Drew inspiration from several experts’ projects and made some tweaks, which significantly improved
- Category
- multimodal media
- Stars
- 1,044
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 96/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 158 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
api-reverse-engineering dreamina image-to-image jimeng text-to-image unofficial-api
Something wrong? Category · Trend · Risk
official code repo for paper "CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers"
- Category
- multimodal media
- Stars
- 955
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1465 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
pretrained-models pytorch text-to-image transformer
Something wrong? Category · Trend · Risk
Implementation of Muse: Text-to-Image Generation via Masked Generative Transformers, in Pytorch
- Category
- multimodal media
- Stars
- 918
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 890 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanisms deep-learning text-to-image transformers
Something wrong? Category · Trend · Risk
[TMLR 2025🔥] A survey for the autoregressive models in vision.
- Category
- multimodal media
- Stars
- 805
- Readiness
- high risk (38/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 16/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 94 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
acceleration autoregressive computer-vision deep-learning diffusion embodied-ai
Something wrong? Category · Trend · Risk
CLIP + FFT/DWT/RGB = text to image/video
- Category
- multimodal media
- Stars
- 790
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 540 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
clip text-to-image text-to-video
Something wrong? Category · Trend · Risk
AI-powered Text-to-Art Generator - Text2Art.com
- Category
- multimodal media
- Stars
- 773
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 1112 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
art colab-notebook deep-learning gan generative-art machine-learning
Something wrong? Category · Trend · Risk
Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
- Category
- multimodal media
- Stars
- 770
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 924 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
A collection of awesome text-to-image generation studies.
- Category
- multimodal media
- Stars
- 763
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence diffusion-models text-to-image text-to-image-ai text-to-image-diffusion
Something wrong? Category · Trend · Risk
Run the official Stable Diffusion releases in a Docker container with txt2img, img2img, depth2img, pix2pix, upscale4x, and inpaint.
- Category
- multimodal media
- Stars
- 746
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 952 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
dall-e dalle diffusion docker generative-art huggingface
Something wrong? Category · Trend · Risk
Personalization for Stable Diffusion via Aesthetic Gradients 🎨
- Category
- multimodal media
- Stars
- 742
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 1386 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models laion stable-diffusion text-to-image text2image
Something wrong? Category · Trend · Risk
A Survey on Text-to-Video Generation/Synthesis.
- Category
- multimodal media
- Stars
- 738
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc image-generation text-to-image text-to-video video-generation
Something wrong? Category · Trend · Risk
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high per
- Category
- multimodal media
- Stars
- 724
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 26/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 154 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc clip controlnet deepseek-vl dit eva-clip
Something wrong? Category · Trend · Risk
The most advanced Nano Banana image generator and editor application. Your central hub for AI image generation and revisions. Intuitive UI features reference images, editing with image masks, version
- Category
- multimodal media
- Stars
- 704
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aiimages bolt imageediting imageeditor imagegeneration imagegenerator
Something wrong? Category · Trend · Risk
HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation
- Category
- multimodal media
- Stars
- 675
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 297 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc diffusion-models diffusion-transformer image-generation text-to-image
Something wrong? Category · Trend · Risk
Official implementation of OneDiffusion paper (CVPR 2025)
- Category
- multimodal media
- Stars
- 662
- Readiness
- high risk (42/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 601 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models image-translation novel-view-synthesis single-view-reconstruction stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Flash Diffusion — accelerating conditional diffusion models (AAAI 2025 Oral)
- Category
- multimodal media
- Stars
- 662
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 514 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models distillation dit inpainting sdxl super-resolution
Something wrong? Category · Trend · Risk
[ICCV 2023] A latent space for stochastic diffusion models
- Category
- multimodal media
- Stars
- 658
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 950 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models generative-models image-synthesis image-to-image-translation score-based-generative-models stable-diffusion
Something wrong? Category · Trend · Risk
基于Stable Diffusion优化的AI绘画模型。支持输入中英文文本,可生成多种现代艺术风格的高质量图像。| An optimized text-to-image model based on Stable Diffusion. Both Chinese and English text inputs are available to generate images. The model c
- Category
- multimodal media
- Stars
- 648
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1247 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-painting aigc artificial-intelligence bert clip cv
Something wrong? Category · Trend · Risk
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
- Category
- multimodal media
- Stars
- 642
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
autoregressive-models generative generative-ai image-gen large-language-models llms
Something wrong? Category · Trend · Risk
face-to-sticker
- Category
- multimodal media
- Stars
- 642
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, documentation. Risks: no push in 889 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai comfyui generative-ai replicate text-to-image
Something wrong? Category · Trend · Risk
Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
- Category
- multimodal media
- Stars
- 632
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 794 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
computer-vision deep-learning diffusion diffusion-models generative-model image-generation
Something wrong? Category · Trend · Risk
[CVPR 2024 Highlight] MIGC and [TPAMI 2024] MIGC++ (Official Implementation)
- Category
- multimodal media
- Stars
- 612
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 449 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc computer-vision cvpr cvpr2024 stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Generative Adversarial Text to Image Synthesis / Please Star -->
- Category
- multimodal media
- Stars
- 599
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 2023 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
gan tensorflow tensorlayer text-to-image
Something wrong? Category · Trend · Risk
Official code for the CVPR 2025 paper "SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models."
- Category
- multimodal media
- Stars
- 589
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 432 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cvpr2025 diffusion-models drawing huggingface-spaces image-generation stable-diffusion
Something wrong? Category · Trend · Risk
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
- Category
- multimodal media
- Stars
- 588
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 794 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
blended-diffusion deep-learning diffusion multimodal openai openai-clip
Something wrong? Category · Trend · Risk
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
- Category
- multimodal media
- Stars
- 551
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 20/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 490 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc large-language-models large-vision-language-models llm lvlm mllm
Something wrong? Category · Trend · Risk
T2F: text to face generation using Deep Learning
- Category
- multimodal media
- Stars
- 546
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1546 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
adversarial-machine-learning gan generative-adversarial-network progressively-growing-gan text-to-image
Something wrong? Category · Trend · Risk
Implementation of Parti, Google's pure attention-based text-to-image neural network, in Pytorch
- Category
- multimodal media
- Stars
- 537
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 973 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanism deep-learning text-to-image transformers
Something wrong? Category · Trend · Risk
Multimodal AI Story Teller, built with Stable Diffusion, GPT, and neural text-to-speech
- Category
- multimodal media
- Stars
- 535
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1074 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ddpm diffusion-models gpt image-generation natural-language-generation pytorch
Something wrong? Category · Trend · Risk
[NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
- Category
- multimodal media
- Stars
- 530
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 266 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
autoregressive-models generative generative-ai generative-model image-generation image-tokenizer
Something wrong? Category · Trend · Risk
Official implementation for "Break-A-Scene: Extracting Multiple Concepts from a Single Image" [SIGGRAPH Asia 2023]
- Category
- multimodal media
- Stars
- 525
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 936 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusion-models generative-ai multimodal text-to-image
Something wrong? Category · Trend · Risk
StyleShot: A SnapShot on Any Style. 一款可以迁移任意风格到任意内容的模型,无需针对图片微调,即能生成高质量的个性风格化图片!
- Category
- multimodal media
- Stars
- 472
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 403 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
controllable-generation style-transfer text-to-image
Something wrong? Category · Trend · Risk
[ ICLR 2024 ] Official Codebase for "InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists"
- Category
- multimodal media
- Stars
- 461
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 832 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models generative-model multi-task-learning stable-diffusion text-to-image vision-language-model
Something wrong? Category · Trend · Risk
A CLI tool/python module for generating images from text using guided diffusion and CLIP from OpenAI.
- Category
- multimodal media
- Stars
- 459
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 18/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 220 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning diffusion image-generation multimodal multimodality
Something wrong? Category · Trend · Risk
🎨 精选 3000+ Gemini Nano Banana Pro 高质量提示词与生成案例 | 涵盖摄影、设计、艺术、营销等多领域 | 双语支持 | JSON 格式
- Category
- multimodal media
- Stars
- 454
- Readiness
- high risk (41/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art-prompt-engineering ai-image-generation ai-prompts awesome awesome-list creative-prompts
Something wrong? Category · Trend · Risk
An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.
- Category
- multimodal media
- Stars
- 451
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 248 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-editing multimodal-large-language-models text-to-image
Something wrong? Category · Trend · Risk
Open-AI's DALL-E for large scale training in mesh-tensorflow.
- Category
- multimodal media
- Stars
- 432
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1637 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence autoregressive multimodal text-to-image transformers variational-autoencoder
Something wrong? Category · Trend · Risk
🧠 世界上覆盖最全的优秀Qwen提示语大全,欢迎贡献你的提示词。🧠 The most comprehensive collection of excellent Qwen prompts in the world. Feel free to contribute your own prompts!
- Category
- multimodal media
- Stars
- 412
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome claude example gpt openai prompt
Something wrong? Category · Trend · Risk
Pytorch implementation of Generative Adversarial Text-to-Image Synthesis paper
- Category
- multimodal media
- Stars
- 411
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 2205 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
gans image-generation pytorch text-to-image zero-shot-learning
Something wrong? Category · Trend · Risk
[NeurIPS'23] "MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing".
- Category
- multimodal media
- Stars
- 411
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 533 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models image-editing image-generation image-synthesis instruction-following text-to-image
Something wrong? Category · Trend · Risk
Official implementation for "Stable Flow: Vital Layers for Training-Free Image Editing" [CVPR 2025]
- Category
- multimodal media
- Stars
- 409
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 425 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusion-models flow-models generative-ai image-editing machine-learning
Something wrong? Category · Trend · Risk
Officail Implementation for "Cross-Image Attention for Zero-Shot Appearance Transfer"
- Category
- multimodal media
- Stars
- 404
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 824 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
appearance-transfer diffusion-models stable-diffusion style-transfer text-to-image
Something wrong? Category · Trend · Risk
📚 Collection of awesome generation acceleration resources.
- Category
- multimodal media
- Stars
- 402
- Readiness
- high risk (38/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 16/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 396 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models efficient-deep-learning efficient-inference image-generation model-acceleration text-to-image
Something wrong? Category · Trend · Risk
End-to-end recipes for optimizing diffusion models with torchao and diffusers (inference and FP8 training).
- Category
- multimodal media
- Stars
- 400
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 211 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
architecture-optimization diffusion-models flux text-to-image torch torch-compile
Something wrong? Category · Trend · Risk
Just playing with getting CLIP Guided Diffusion running locally, rather than having to use colab.
- Category
- multimodal media
- Stars
- 385
- Readiness
- high risk (42/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 1439 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
openai-clip text-to-image text2image
Something wrong? Category · Trend · Risk
Official GitHub repository for FLUX.1 Krea [dev].
- Category
- multimodal media
- Stars
- 365
- Readiness
- high risk (39/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 371 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models flux machine-learning text-to-image
Something wrong? Category · Trend · Risk
A list of AI Art courses, tools, libraries, people, and places.
- Category
- multimodal media
- Stars
- 364
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 777 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-art artificial-intelligence awesome awesome-list creative-coding deep-learning
Something wrong? Category · Trend · Risk
Official Repository of the paper "Trajectory Consistency Distillation"
- Category
- multimodal media
- Stars
- 360
- Readiness
- high risk (38/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 831 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
consistency-models diffusion fast-sampling score-based-models stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
[Neurips 2023 & TPAMI] T2I-CompBench (++) for Compositional Text-to-image Generation Evaluation
- Category
- multimodal media
- Stars
- 345
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 15/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 92 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
benchmark compositionality dataset text-to-image
Something wrong? Category · Trend · Risk
This is a ChatGPT based prompt generation model for MidJorney. The purpose of this model is to simplify the creation of images and increase their creativity. By introducing a partial hint, ChatGPT cre
- Category
- multimodal media
- Stars
- 341
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-art ai-painting chatgpt midjourney prompt
Something wrong? Category · Trend · Risk
FIBO is a SOTA, first open-source, JSON-native text-to-image model built for controllable, predictable, and legally safe image generation.
- Category
- multimodal media
- Stars
- 329
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 212 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai ai controlable-image-generation creative-ai deep-learning enterprise-ready
Something wrong? Category · Trend · Risk
[CVPR2022 oral] A Simple and Effective Baseline for Text-to-Image Synthesis
- Category
- multimodal media
- Stars
- 326
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest. Risks: no push in 317 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
generative-adversarial-network text-to-image
Something wrong? Category · Trend · Risk
AlignProp uses direct reward backpropogation for the alignment of large-scale text-to-image diffusion models. Our method is 25x more sample and compute efficient than reinforcement learning methods (P
- Category
- multimodal media
- Stars
- 324
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 645 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alignment diffusion-models reinforcement-learning stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Implementation of Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models
- Category
- multimodal media
- Stars
- 323
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1202 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusion-models stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
Colab notebook for Stable Diffusion Hyper-SDXL.
- Category
- multimodal media
- Stars
- 323
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 27/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 479 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
colab colab-notebook colaboratory deep-learning diffusers diffusion
Something wrong? Category · Trend · Risk
Ovis-Image is a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints.
- Category
- multimodal media
- Stars
- 321
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-generation text-to-image
Something wrong? Category · Trend · Risk
🔥ICLR 2025 (Spotlight) One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
- Category
- multimodal media
- Stars
- 320
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 291 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion diffusion-model diffusion-models storytelling text-to-image
Something wrong? Category · Trend · Risk
Getting the latest versions of Disco Diffusion to work locally, instead of colab. Including how I run this on Windows, despite some Linux only dependencies ;)
- Category
- multimodal media
- Stars
- 316
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1472 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
3d-animation art clip dall-e dalle disco-diffusion
Something wrong? Category · Trend · Risk
MinImagen: A minimal implementation of the Imagen text-to-image model
- Category
- multimodal media
- Stars
- 312
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 1187 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusion-models imagen pytorch super-resolution text-to-image
Something wrong? Category · Trend · Risk
🔥 [ICCV 2025 Highlight] Official ComfyUI native node supporting InfiniteYou with FLUX
- Category
- multimodal media
- Stars
- 298
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 378 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
comfyui comfyui-nodes diffusion diffusion-transformer dit face
Something wrong? Category · Trend · Risk
[ICML2025] An 8-step inversion and 8-step editing process works effectively with the FLUX-dev model. (3x speedup with results that are comparable or even superior to baseline methods)
- Category
- multimodal media
- Stars
- 295
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 463 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
flux flux-dev generative-model image-editing text-to-image
Something wrong? Category · Trend · Risk
Helios: Real Real-Time Long Video Generation Model
- Category
- multimodal media
- Stars
- 2,034
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
acceleration diffusion diffusion-model diffusion-models efficient-tuning high-quality
Something wrong? Category · Trend · Risk
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
- Category
- multimodal media
- Stars
- 1,166
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +320 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-video claude-code claude-skill collage-video explainer-video
Something wrong? Category · Trend · Risk
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
- Category
- multimodal media
- Stars
- 911
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auto-regressive-diffusion-model autoregressive-models consistency-models diffusion diffusion-models distillation
Something wrong? Category · Trend · Risk
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
- Category
- multimodal media
- Stars
- 521
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agents ai-video cli coding-agents ffmpeg generative-ai
Something wrong? Category · Trend · Risk
AI short drama & micro-drama video generator — turns any idea into a complete short-form drama using multi-agent AI pipeline (screenwriter → storyboard → frames → video). Seedance 2 VIP, Kling 3.0 Pro
- Category
- multimodal media
- Stars
- 451
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +14 stars in 7 days
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: release recency.
agentic-ai ai-drama ai-filmmaking ai-short-drama ai-video-generation drama-generator
Something wrong? Category · Trend · Risk
A curated archive of breakthroughs in Agents, Architecture, Training, RAG, and On-Device AI.
- Category
- multimodal media
- Stars
- 421
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- watch
- Maintenance risk
- 0/100 · high confidence
Why: +2 stars in 7 days
Why it may be a gem: open issue backlog is stable or shrinking
Strongest signals: push recency, issue load, documentation. Risks: maintenance is concentrated in one contributor. Missing inputs: release recency.
all-to-all dialogue distillation efficient llm multimodal
Something wrong? Category · Trend · Risk
Native macOS app for AI video generation using LTX-Video model, optimized for Apple Silicon
- Category
- multimodal media
- Stars
- 371
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-video ltx-2 ltx-video mac-app mac-native text-to-video
Something wrong? Category · Trend · Risk
AI video generation SDK — JSX for videos. One API for Kling, Flux, ElevenLabs, Veed. Built on Vercel AI SDK.
- Category
- multimodal media
- Stars
- 333
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-sdk ai-video claude-code cursor elevenlabs flux
Something wrong? Category · Trend · Risk
Claude Code Skill that turns any idea into a cinematic, model-ready video prompt — Sora · Kling · Veo · Seedance. 21 genre templates, 5-stage structure, eval-tested. Distilled from the AI short Hollyw
- Category
- multimodal media
- Stars
- 328
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video chinese cinematic cinematic-prompts claude-code claude-skill
Something wrong? Category · Trend · Risk
Clip any moment from any video with prompts
- Category
- multimodal media
- Stars
- 285
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video ai-video-generator ai-video-maker clipanything clipbasic shorts
Something wrong? Category · Trend · Risk
An all-in-one, 100% local AI video, image, and music studio. Director mode plans full music videos and short films from a single prompt. Built on the WanGP pipeline. Install via Pinokio.
- Category
- multimodal media
- Stars
- 239
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +137 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-video image-generation local-ai ltx-video music-generation
Something wrong? Category · Trend · Risk
Open AI UGC — free, open-source alternative to Arcads and MakeUGC. Generate AI UGC video ads with realistic AI actors using Veo 3.1, Seedance 2, Grok Video, and Happy Horse 1. Self-host on Next.js + S
- Category
- multimodal media
- Stars
- 233
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-actors ai-marketing ai-ugc ai-video-ads ai-video-generator arcads-alternative
Something wrong? Category · Trend · Risk
Open-source, self-hosted AI video generator — completely free. Text to multi-scene video with narration, subtitles, and digital anchor via Web UI, powered by Agnes AI.
- Category
- multimodal media
- Stars
- 214
- Readiness
- ready (99/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +17 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agnes agnes-video ai-video-generator ai-video-maker free-ai-software free-ai-video
Something wrong? Category · Trend · Risk
自动生成电商带货短视频:上传商品图,AI 生成脚本、素材并合成视频,适配抖音、快手、小红书
- Category
- multimodal media
- Stars
- 205
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +13 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video-generation ffmpeg flux kling kuaishou nextjs
Something wrong? Category · Trend · Risk
🦦 Crayotter: A Multimodal AI-Agent for Video-Editing, Video-Composing, and Video Production. Powered by Multimodal LLMs for autonomous Text-to-Video agentic framework. | 基于多模态大模型 (Multimodal LLMs) 的 A
- Category
- multimodal media
- Stars
- 189
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent mllms text-to-video video-composing video-editing video-production
Something wrong? Category · Trend · Risk
The most complete, up-to-date comparison of AI video generation models — which model, via which API, at what price, and how fast.
- Category
- multimodal media
- Stars
- 179
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video ai-video-generator awesome awesome-list generative-ai hailuo
Something wrong? Category · Trend · Risk
Frontier AI video generation for coding agents - Seedance, Runway, Veo, Kling on your own keys, behind a hard cost cap. CLI + MCP.
- Category
- multimodal media
- Stars
- 164
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-agents claude cli developer-tools ffmpeg
Something wrong? Category · Trend · Risk
Best Open Source AI Video Generator 2026: Text to Multi-Scene Movies with Narration & Subtitles
- Category
- multimodal media
- Stars
- 157
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agnes agnes-video ai-video-generator ai-video-maker free-ai-software free-ai-video
Something wrong? Category · Trend · Risk
2026 Guide to Self-Hosted Open Source AI Video Generation
- Category
- multimodal media
- Stars
- 156
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agnes agnes-video ai-video-generator ai-video-maker free-ai-software free-ai-video
Something wrong? Category · Trend · Risk
A timeline editor for MiniMax H3 inside ComfyUI - storyboard prompts, first/last keyframes, image/video/audio references, joint audio, live sampling preview, retakes and shot chaining.
- Category
- multimodal media
- Stars
- 153
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
comfyui comfyui-nodes image-to-video minimax minimax-h3 storyboard
Something wrong? Category · Trend · Risk
🎬 Turn any topic into a finished Vox-style paper-collage explainer / motion graphics video — script, collage keyframes, animation, voice-over, music & captions, all automated. An agent skill for Claud
- Category
- multimodal media
- Stars
- 138
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +23 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skill ai-motion-graphics ai-video claude-code collage explainer-video
Something wrong? Category · Trend · Risk
剧本分镜工具智能体(PenShot):电影/动漫/短剧/小说/剧本→分镜→片段→prompt | 基于 LangGraph+LLM,自动解析任意格式剧本,生成 Sora/Veo/Runway 等模型可用的连贯text-to-video提示词。保持角色/剧情跨片段一致,支持 MCP/REST API/函数调用 | Python库 + A2A集成。(LLM-powered screenplay-to-
- Category
- multimodal media
- Stars
- 134
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-to-agent ai-filmmaking ai-video-generation character-consistency function-calling kling-ai
Something wrong? Category · Trend · Risk
100+ curated Seedance 2.5 prompts with real video previews, plus an installable Agent Skill that optimizes prompts, creates storyboards, and generates videos via Seedance models.
- Category
- multimodal media
- Stars
- 132
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +22 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills ai ai-agents ai-video artificial-intelligence awesome-list
Something wrong? Category · Trend · Risk
Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.
- Category
- multimodal media
- Stars
- 121
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent agentic-ai ai ai-audio ai-video claude
Something wrong? Category · Trend · Risk
AI 视频制作与自媒体运营工具精选合集 — 涵盖一键生成、文生视频、多平台分发、字幕翻译等 150+ 开源项目,每周更新。
- Category
- multimodal media
- Stars
- 113
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video automation awesome awesome-list content-creation short-video
Something wrong? Category · Trend · Risk
Open source AI video engine built on Remotion. Voice cloning, AI avatars, animated captions, text-to-video (Wan 2.2/LTX), AI music, video editor, timeline, 100+ transitions. Free Synthesia, HeyGen, Ru
- Category
- multimodal media
- Stars
- 95
- Readiness
- ready (79/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-avatars ai-music ai-video auto-captions background-removal capcut-alternative
Something wrong? Category · Trend · Risk
Seedance 2.5 Shot Design Skills
- Category
- multimodal media
- Stars
- 90
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video claude-skills jimeng prompt-engineering seedance text-to-video
Something wrong? Category · Trend · Risk
颠覆性AI内容再创作引擎,通过“解构-重构”爆款视频模式,全自动生成高度原创短视频。AIGC, Content Creation, Video Generation, Automation, LLM, Python, FFmpeg.
- Category
- multimodal media
- Stars
- 90
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc automation content-creation content-repurposing content-strategy ffmpeg
Something wrong? Category · Trend · Risk
Agent-native AI video factory: one brief → narrated, subtitled, scene-editable videos (shorts & long-form). Deterministic HTML rendering, free keyless stack, Korean-first.
- Category
- multimodal media
- Stars
- 76
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video-generator faceless-video hyperframes shorts-automation text-to-video video-agent
Something wrong? Category · Trend · Risk
Open-source Next.js SaaS for Seedance 2.0 , Seedance 2.5 and Seedance 2 Mini video generation — Stripe billing, credits, NextAuth, and Prisma out of the box.
- Category
- multimodal media
- Stars
- 72
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-saas ai-video ai-video-generator bytedance-seedance generative-ai generative-video
Something wrong? Category · Trend · Risk
MangaV (漫织 AI) - 端到端 AI 漫剧创作平台。集成 13 大 AI 大模型、视听多模态与 FFmpeg 硬件加速引擎,输入一本小说,AI 自动生成 4K 精致漫剧。
- Category
- multimodal media
- Stars
- 72
- Readiness
- ready (100/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-manga ai-tools anime-generator computer-vision cyberpunk
Something wrong? Category · Trend · Risk
Biến 1 URL tin tức/GitHub thành video 9:16 chuẩn TikTok/Reels/Shorts trong 5 phút — không cần edit, TTS tiếng Việt + motion graphic tự động.
- Category
- multimodal media
- Stars
- 45
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video-generation shorts-generator text-to-video tiktok-video
Something wrong? Category · Trend · Risk
Python SDK for the MiniMax H3 API — text-to-video, image-to-video, and first/last-frame video generation via Muapi.
- Category
- multimodal media
- Stars
- 45
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +16 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-tools ai-video api artificial-intelligence creative-ai developer-tools
Something wrong? Category · Trend · Risk
Ready-to-run AWS Deadline Cloud samples for rendering (Maya, Blender, Houdini, Nuke); physical AI (robotics, autonomous driving simulation, MuJoCo, CARLA); synthetic data generation; generative AI vid
- Category
- multimodal media
- Stars
- 40
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aws aws-deadline-cloud bioinformatics blender carla cloud-rendering
Something wrong? Category · Trend · Risk
Seedance 2.0, 2.5 & 2 Mini ComfyUI: custom nodes and workflows for ByteDance video generation via MuAPI — text-to-video, image-to-video, Omni Reference, consistent characters, and video extend.
- Category
- multimodal media
- Stars
- 38
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-video ai-video-generation bytedance comfyui comfyui-custom-nodes
Something wrong? Category · Trend · Risk
Compare and generate AI videos across Sora, Veo, Kling, Seedance & more.
- Category
- multimodal media
- Stars
- 35
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-tools ai-video-generation ai-video-generator fal-ai generative-ai image-to-video
Something wrong? Category · Trend · Risk
🎬 Curated Grok Imagine (xAI) video generation prompts — cinematic, action, anime, product, meme styles. Includes prompt engineering tips, style guides, and creative video workflows.
- Category
- multimodal media
- Stars
- 34
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-video awesome grok grok-imagine prompts
Something wrong? Category · Trend · Risk
A curated and auto-updated collection of video diffusion / video generation papers from arXiv, covering text-to-video, image-to-video, controllable generation, world models, video editing, and 16+ res
- Category
- multimodal media
- Stars
- 34
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
arxiv arxiv-papers awesome awesome-list controllable-generation diffusion-models
Something wrong? Category · Trend · Risk
NEWTON: Agentic Planning for Physically Grounded Video Generation
- Category
- multimodal media
- Stars
- 32
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic-ai agentic-video-generation ai-agents aigc computer-vision diffusion-models
Something wrong? Category · Trend · Risk
Claude skill that turns a single request into a narrated mythic short film in a layered cut-paper diorama style — three-colour cardstock worlds, paper wipes, Unsora MCP + local ffmpeg, one assembled M
- Category
- multimodal media
- Stars
- 31
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video animation claude claude-code claude-skill claude-skills
Something wrong? Category · Trend · Risk
AI 驱动的短视频制作系统。给它一个想法,它帮你写稿、配音、做画面、出成片。9 阶段 DAG 管线 + 自进化评分,支持每日自动执行。基于 Claude Code + HyperFrames。
- Category
- multimodal media
- Stars
- 29
- Readiness
- ready (96/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation claude-code dag douyin html hyperframes
Something wrong? Category · Trend · Risk
📊 Daily auto-updated snapshots of all Arena AI (LMSYS Chatbot Arena) leaderboards — LLM, Vision, Code, Video, Image & more. Structured JSON with historical tracking.
- Category
- multimodal media
- Stars
- 28
- Readiness
- ready (100/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-benchmark arena-ai benchmark chatbot-arena leaderboard
Something wrong? Category · Trend · Risk
Agentic AI video generator — fully autonomous text-to-video pipeline with free TTS/voice-clone, auto-editing, captions, and 20+ single-task operations. Zero-cost, MIT.
- Category
- multimodal media
- Stars
- 27
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agentic ai ai-video automation captions chatterbox
Something wrong? Category · Trend · Risk
Building an Image-to-Video (I2V) Model from Scratch
- Category
- multimodal media
- Stars
- 25
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
computer-vision deep-learning diffusion-models flow-matching generative-ai huggingface
Something wrong? Category · Trend · Risk
Claude plugin for Kling AI video prompts. Works with Claude.ai, Claude Desktop & Claude Code. I2V, T2V, Motion Control, Storyboards. Try live: maciejdzierzek.com
- Category
- multimodal media
- Stars
- 23
- Readiness
- ready (98/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-tools ai-video claude-code claude-skill generative-ai image-to-video
Something wrong? Category · Trend · Risk
ByteDance Seedance 2.0 / 2.5 from your terminal. Text-to-video, image-to-video, audio & video refs, up to 4K. One agent-ready Rust binary.
- Category
- multimodal media
- Stars
- 23
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-cli agent-friendly ai ai-video bytedance byteplus
Something wrong? Category · Trend · Risk
Helios: Real Real-Time Long Video Generation Model
- Category
- multimodal media
- Stars
- 23
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
acceleration diffusion diffusion-model diffusion-models efficient-tuning high-quality
Something wrong? Category · Trend · Risk
Flux AI Works
- Category
- multimodal media
- Stars
- 21
- Readiness
- ready (73/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
flux-ai flux-api flux-dev image-to-video-api image-to-video-generation text-to-video
Something wrong? Category · Trend · Risk
AI agent skills for video, image, speech & music generation. Works with Claude Code, Cursor, Windsurf, OpenCode, ClawHub.
- Category
- multimodal media
- Stars
- 20
- Readiness
- needs review (68/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent-skills ai-agent ai-tools ai-video claude-code clawhub
Something wrong? Category · Trend · Risk
Sora 2 api access using http://muapi.ai
- Category
- multimodal media
- Stars
- 18
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-video api generative-ai openai openai-api python
Something wrong? Category · Trend · Risk
[NeurIPS 2025] Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
- Category
- multimodal media
- Stars
- 18
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alignment diffusion-models inference-time-scaling neurips-2025 test-time-scaling text-to-video
Something wrong? Category · Trend · Risk
ComfyUI custom nodes for Google Gemini Omni — text-to-video, image-to-video, and video editing via the Gemini Omni API on muapi.ai
- Category
- multimodal media
- Stars
- 18
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: strong signals despite limited visibility; healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-video ai-video-generation comfyui comfyui-custom-nodes comfyui-nodes
Something wrong? Category · Trend · Risk
Official repository for LTX-Video
- Category
- multimodal media
- Stars
- 10,818
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +31 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 214 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models dit image-to-video image-to-video-generation text-to-video text-to-video-generation
Something wrong? Category · Trend · Risk
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
- Category
- multimodal media
- Stars
- 5,071
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 210 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-video text-to-video video-generation
Something wrong? Category · Trend · Risk
HunyuanVideo-1.5: A leading lightweight video generation model
- Category
- multimodal media
- Stars
- 4,556
- Readiness
- high risk (39/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 20/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 119 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-video text-to-video video-generation
Something wrong? Category · Trend · Risk
[CSUR] A Survey on Video Diffusion Models
- Category
- multimodal media
- Stars
- 2,310
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome awesome-list diffusion diffusion-models survey text-to-video
Something wrong? Category · Trend · Risk
Implementation of Make-A-Video, new SOTA text to video generator from Meta AI, in Pytorch
- Category
- multimodal media
- Stars
- 1,985
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 826 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanisms axial-convolutions deep-learning text-to-video
Something wrong? Category · Trend · Risk
AI-powered animated comic generator — transform scripts into fully animated videos with AI-driven character design, storyboarding, and video synthesis.
- Category
- multimodal media
- Stars
- 1,750
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 17/100 · low confidence
Why: +47 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 102 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-animation ai-animation-generator ai-animation-tools aicomicbuilder storytelling storytelling-ai
Something wrong? Category · Trend · Risk
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
- Category
- multimodal media
- Stars
- 1,724
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 23/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 137 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc benchmark dataset evaluation-kit gen-ai stable-diffusion
Something wrong? Category · Trend · Risk
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
- Category
- multimodal media
- Stars
- 1,515
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 330 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc consistency-models text-to-video video-generation
Something wrong? Category · Trend · Risk
Text To Video Synthesis Colab
- Category
- multimodal media
- Stars
- 1,513
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 862 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
colab colab-notebook colaboratory t2v text-to-video
Something wrong? Category · Trend · Risk
Implementation of Video Diffusion Models, Jonathan Ho's new paper extending DDPMs to Video Generation - in Pytorch
- Category
- multimodal media
- Stars
- 1,384
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 826 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence ddpm deep-learning text-to-video video-generation
Something wrong? Category · Trend · Risk
[TPAMI 2025🔥] MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
- Category
- multimodal media
- Stars
- 1,339
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 19/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 115 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models long-video-generation metamorphic-video-generation open-sora-plan text-to-video time-lapse
Something wrong? Category · Trend · Risk
✨ Hotshot-XL: State-of-the-art AI text-to-GIF model trained to work alongside Stable Diffusion XL
- Category
- multimodal media
- Stars
- 1,111
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 927 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai hotshot hotshot-xl sdxl text-to-gif text-to-video
Something wrong? Category · Trend · Risk
[NeurIPS 2024] An official implementation of "ShareGPT4Video: Improving Video Understanding and Generation with Better Captions"
- Category
- multimodal media
- Stars
- 1,092
- Readiness
- high risk (39/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 667 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chatgpt gpt gpt-4v large-language-models large-multimodal-models large-video-language-models
Something wrong? Category · Trend · Risk
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.
- Category
- multimodal media
- Stars
- 1,060
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 313 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc-audio foley-art foley-sound-synthesis text-to-audio text-to-video text-video-to-audio
Something wrong? Category · Trend · Risk
[ECCV 2024 Oral] MotionDirector: Motion Customization of Text-to-Video Diffusion Models.
- Category
- multimodal media
- Stars
- 1,049
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 716 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models motion-customization text-to-motion text-to-video text-to-video-generation video-generation
Something wrong? Category · Trend · Risk
Text to video generator in the brainrot form. Learn about any topic from your favorite personalities 😼.
- Category
- multimodal media
- Stars
- 955
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 17/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 104 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automate chatgpt nextjs python remotion text-to-video
Something wrong? Category · Trend · Risk
Video generation from text&image, 1st-gen
- Category
- multimodal media
- Stars
- 921
- Readiness
- high risk (35/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 454 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-video text-to-video video-generation
Something wrong? Category · Trend · Risk
[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
- Category
- multimodal media
- Stars
- 851
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 19/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 115 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion diffusion-models identity-preserving text-to-video video-generation video-generation-dataset
Something wrong? Category · Trend · Risk
Kandinsky 5.0: A family of diffusion models for Video & Image generation
- Category
- multimodal media
- Stars
- 806
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion distillation kandinsky text-to-video video video-generation
Something wrong? Category · Trend · Risk
Implementation of Phenaki Video, which uses Mask GIT to produce text guided videos of up to 2 minutes in length, in Pytorch
- Category
- multimodal media
- Stars
- 790
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 739 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanisms deep-learning imagination-machine text-to-video transformers
Something wrong? Category · Trend · Risk
Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
- Category
- multimodal media
- Stars
- 777
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 255 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
3d-generation 3d-reconstruction 3d-scene-generation aigc aigc3d genie
Something wrong? Category · Trend · Risk
Finetune ModelScope's Text To Video model using Diffusers 🧨
- Category
- multimodal media
- Stars
- 703
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 967 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning diffusers diffusion-models modelscope pytorch stable-diffusion
Something wrong? Category · Trend · Risk
The official code of Yume
- Category
- multimodal media
- Stars
- 681
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 205 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-to-video interactive-generation real-time-generation text-to-video world-model
Something wrong? Category · Trend · Risk
Multi-model DAG-driven parallel AI film generation — parallel speedup scales with scene independence; Generate film scenes simultaneously instead of one by one; "把影视生成的执行图从拓扑序变成关键路径最优调度" ; 唯一把场景叙事依赖建
- Category
- multimodal media
- Stars
- 641
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 22/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 134 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-cinematography ai-filmmaking ai-video-generation cpm critical-path-method
Something wrong? Category · Trend · Risk
Official codes of VEnhancer: Generative Space-Time Enhancement for Video Generation
- Category
- multimodal media
- Stars
- 578
- Readiness
- high risk (39/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 690 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc-enhancement diffusion-models frame-interpolation text-to-video video-enhancement video-generation
Something wrong? Category · Trend · Risk
Let's finetune video generation models!
- Category
- multimodal media
- Stars
- 551
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, documentation. Risks: no push in 326 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai aigc content-production fine-tuning-diffusion text-to-video video-generation
Something wrong? Category · Trend · Risk
Implementation of NÜWA, state of the art attention network for text to video synthesis, in Pytorch
- Category
- multimodal media
- Stars
- 548
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1298 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence attention-mechanism deep-learning text-to-audio text-to-video transformers
Something wrong? Category · Trend · Risk
[ECCV 2024] FreeInit: Bridging Initialization Gap in Video Diffusion Models
- Category
- multimodal media
- Stars
- 544
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 932 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc text-to-video video-diffusion-model video-generation
Something wrong? Category · Trend · Risk
[AAAI-2026]FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
- Category
- multimodal media
- Stars
- 484
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +24 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 520 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models efficient-generative-model generative-models text-to-video video-generation
Something wrong? Category · Trend · Risk
Codes for ID-Specific Video Customized Diffusion
- Category
- multimodal media
- Stars
- 458
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 897 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models image-animation image-editting personalized-generation text-to-video video-diffusion
Something wrong? Category · Trend · Risk
ICASSP 2022: "Text2Video: text-driven talking-head video synthesis with phonetic dictionary".
- Category
- multimodal media
- Stars
- 438
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 1161 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc avatar deep-learning digital-humanities gan generative-ai
Something wrong? Category · Trend · Risk
A curated collection of AI tools, utilities, and resources for developers and creators
- Category
- multimodal media
- Stars
- 436
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai ai-agent-directory ai-agents-framework ai-artificial-intelligence ai-tools-directory chatbot
Something wrong? Category · Trend · Risk
Code for "Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text" (NeurIPS 2024).
- Category
- multimodal media
- Stars
- 381
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 511 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
3d 3d-aigc 3dgs aigc generative-model text-to-3d
Something wrong? Category · Trend · Risk
Code repository for T2V-Turbo and T2V-Turbo-v2
- Category
- multimodal media
- Stars
- 312
- Readiness
- high risk (36/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load. Risks: no push in 553 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
consistency-models text-to-video video-generation
Something wrong? Category · Trend · Risk
The official implementation for "Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising".
- Category
- multimodal media
- Stars
- 308
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 292 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models long-video-generation stable-diffusion text-to-video text2video video-editing
Something wrong? Category · Trend · Risk
🎥 Create youtube videos from a text prompt in seconds
- Category
- multimodal media
- Stars
- 303
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 999 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation content-creation text-to-video youtube youtube-automation
Something wrong? Category · Trend · Risk
Official PyTorch implementation of TATS: A Long Video Generation Framework with Time-Agnostic VQGAN and Time-Sensitive Transformer (ECCV 2022)
- Category
- multimodal media
- Stars
- 288
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 828 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-to-video long-video-generation pytorch text-to-video video-generation video-manipulation
Something wrong? Category · Trend · Risk
[CVPR 2024] | LAMP: Learn a Motion Pattern for Few-Shot Based Video Generation
- Category
- multimodal media
- Stars
- 283
- Readiness
- high risk (42/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 837 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc diffusion diffusion-model diffusion-models few-shot-learning stable-diffusion
Something wrong? Category · Trend · Risk
Implementation of Lumiere, SOTA text-to-video generation from Google Deepmind, in Pytorch
- Category
- multimodal media
- Stars
- 282
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 742 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence deep-learning denoising-diffusion text-to-video
Something wrong? Category · Trend · Risk
VideoGen-Eval: Agent-based System for Video Generation Evaluation
- Category
- multimodal media
- Stars
- 269
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 234 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc benchmark image-to-video sora-video-ai text-to-video video-evaluation
Something wrong? Category · Trend · Risk
Open-source clone of the MidJourney web interface featuring real AI image and video generation powered by Google's Gemini SDK. Use Imagen 4 to generate images and Veo 2 and 3 for image and text to vid
- Category
- multimodal media
- Stars
- 263
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 382 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai generative-ai image-generation image-to-video open-source text-to-image
Something wrong? Category · Trend · Risk
Avatar Generation For Characters and Game Assets Using Deep Fakes
- Category
- multimodal media
- Stars
- 234
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 719 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
api audio-to-video avatar-generation conversational-ai deep-fake elevenlabs
Something wrong? Category · Trend · Risk
In this blog, we will build a small scale text-to-video model from scratch. We will input a text prompt, and our trained model will generate a video based on that prompt.
- Category
- multimodal media
- Stars
- 232
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 775 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai deep-learning gan gemini gpt-4 openai
Something wrong? Category · Trend · Risk
[NeurIPS 2025 D&B🔥] OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
- Category
- multimodal media
- Stars
- 226
- Readiness
- needs review (50/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc benchmark dataset evaluation-kit image-to-video image-to-video-generation
Something wrong? Category · Trend · Risk
[NeurIPS 2024] AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
- Category
- multimodal media
- Stars
- 215
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, documentation, license. Risks: no push in 314 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
diffusion-models distributed-computing efficient-inference inference-acceleration stable-diffusion text-to-image
Something wrong? Category · Trend · Risk
[NeurIPS 2024 D&B Spotlight🔥] ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
- Category
- multimodal media
- Stars
- 213
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 19/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, documentation, license. Risks: no push in 115 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aigc benchmark dataset diffusion-models evaluation evaluation-kit
Something wrong? Category · Trend · Risk
Official implementations for paper: LivePhoto: Real Image Animation with Text-guided Motion Control
- Category
- multimodal media
- Stars
- 204
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals; strong signals despite limited visibility
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-animation image-to-video text-to-video video-generation
Something wrong? Category · Trend · Risk
Port of OpenAI's Whisper model in C/C++
- Category
- multimodal media
- Stars
- 52,674
- Readiness
- ready (94/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: +203 stars in 7 days; 100+ commits in 30 days
Why it may be a gem: consistent human and community activity; healthy maintenance and project fundamentals
Strongest signals: push recency, commit activity, contributor breadth. Risks: None identified. Missing inputs: None.
Capped lower bounds: 30-day commits, lifetime contributors.
inference openai speech-recognition speech-to-text transformer whisper
Something wrong? Category · Trend · Risk
A free, open source, and extensible speech-to-text application that works completely offline.
- Category
- multimodal media
- Stars
- 28,978
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +746 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility cross-platform speech-to-text tauri-v2
Something wrong? Category · Trend · Risk
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
- Category
- multimodal media
- Stars
- 23,473
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +110 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr speech speech-recognition speech-to-text whisper
Something wrong? Category · Trend · Risk
Translate the video from one language to another and embed dubbing & subtitles.
- Category
- multimodal media
- Stars
- 18,620
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +111 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech-to-text text-to-speech video-transition
Something wrong? Category · Trend · Risk
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android
- Category
- multimodal media
- Stars
- 14,033
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +123 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aarch64 android arm32 asr cpp csharp
Something wrong? Category · Trend · Risk
Build local voice agents with open-source models
- Category
- multimodal media
- Stars
- 11,584
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1,828 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai assistant language-model machine-learning python speech
Something wrong? Category · Trend · Risk
Speech recognition module for Python, supporting several engines and APIs, online and offline.
- Category
- multimodal media
- Stars
- 8,982
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio python speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
💬 Speech recognition for your site
- Category
- multimodal media
- Stars
- 6,818
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech speech-recognition speech-to-text voice
Something wrong? Category · Trend · Risk
Silero Models: pre-trained text-to-speech models made embarrassingly simple
- Category
- multimodal media
- Stars
- 6,054
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +21 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
armenian azerbaijani belarus colab georgian kazakh
Something wrong? Category · Trend · Risk
Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.
- Category
- multimodal media
- Stars
- 5,260
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +195 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai anthropic cross-platform gemini groq linux
Something wrong? Category · Trend · Risk
开源 AI 视频本地化工具:自动完成 YouTube/Bilibili 视频下载、字幕识别与翻译、语音克隆配音、音轨混合和字幕压制。
- Category
- multimodal media
- Stars
- 5,250
- Readiness
- ready (77/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +36 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-dubbing bilibili demucs fastapi ffmpeg nextjs
Something wrong? Category · Trend · Risk
On-device subtitle generation that connects directly to DaVinci Resolve, Premiere, and After Effects.
- Category
- multimodal media
- Stars
- 3,982
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +38 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai cross-platform davinci davinci-resolve premiere resolve
Something wrong? Category · Trend · Risk
Lightweight and powerful real-time audio/speech translation tool based on Windows LiveCaptions.
- Category
- multimodal media
- Stars
- 3,466
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +41 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
api api-integration audio-to-text livecaptions real-time speech-to-text
Something wrong? Category · Trend · Risk
Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative
- Category
- multimodal media
- Stars
- 3,086
- Readiness
- ready (87/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alexa deep-learning echo esp-adf esp-idf esp32
Something wrong? Category · Trend · Risk
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
- Category
- multimodal media
- Stars
- 3,052
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai audio-to-text golang speech-recognition speech-to-text stt
Something wrong? Category · Trend · Risk
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
- Category
- multimodal media
- Stars
- 2,628
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +55 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ane asr audio automatic-speech-recognition avfoundation coreml
Something wrong? Category · Trend · Risk
No description
- Category
- multimodal media
- Stars
- 2,497
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-sdk chat-completions foundry-local gpu-acceleration local-ai microsoft
Something wrong? Category · Trend · Risk
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for n
- Category
- multimodal media
- Stars
- 2,377
- Readiness
- ready (74/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +21 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-assistant claude cluely-alternative desktop-app gemini grok
Something wrong? Category · Trend · Risk
ggml speech-to-text inference for 16+ model families
- Category
- multimodal media
- Stars
- 1,712
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +58 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr ggml gguf speech-to-text
Something wrong? Category · Trend · Risk
Local speech-to-text for macOS on-device AI, fully private, optional cloud
- Category
- multimodal media
- Stars
- 1,675
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +21 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon dictation macos on-device privacy speech-to-text
Something wrong? Category · Trend · Risk
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
- Category
- multimodal media
- Stars
- 1,557
- Readiness
- needs review (71/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr flatpak-applications linux-desktop machine-translation nmt offline
Something wrong? Category · Trend · Risk
Custom nodes that extend the capabilities of Comfyui
- Category
- multimodal media
- Stars
- 1,520
- Readiness
- ready (84/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
glm ide painter pose-detection speech-recognition speech-synthesis
Something wrong? Category · Trend · Risk
Native speech-to-text for Linux - Fast, accurate, private, and hackable system-wide dictation
- Category
- multimodal media
- Stars
- 1,136
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +14 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai archlinux cachyos cohere-ai debian dictation
Something wrong? Category · Trend · Risk
Voice-to-text with push-to-talk for Wayland compositors
- Category
- multimodal media
- Stars
- 1,050
- Readiness
- ready (89/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility dictation gnome hyprland kde linux
Something wrong? Category · Trend · Risk
Open-source voice-AI SDK. The Vapi/Retell alternative for builders who want to own the stack. Give your AI agent a phone number in 4 lines — Python and TypeScript, MIT licensed, Twilio, Telnyx, and Pl
- Category
- multimodal media
- Stars
- 1,026
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-agent ai-phone-agent hermes-agent llm mastra open-source
Something wrong? Category · Trend · Risk
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
- Category
- multimodal media
- Stars
- 1,010
- Readiness
- ready (97/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automatic-speech-recognition conformer contextnet ctc deepspeech2 end2end
Something wrong? Category · Trend · Risk
Open source voice dictation technology
- Category
- multimodal media
- Stars
- 998
- Readiness
- ready (80/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
agent agentic-ai dictation free local-ai open-source
Something wrong? Category · Trend · Risk
Fully local, private and cross platform Speech-to-Text with LLM Post-processing
- Category
- multimodal media
- Stars
- 990
- Readiness
- ready (82/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr asr-model debian-packages linux macos privacy
Something wrong? Category · Trend · Risk
Whisper.net. Speech to text made simple using Whisper Models
- Category
- multimodal media
- Stars
- 940
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cross-platform dotnet dotnetcore speech-recognition speech-to-text translation
Something wrong? Category · Trend · Risk
.NET MAUI Samples
- Category
- multimodal media
- Stars
- 915
- Readiness
- ready (85/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
azure blazor dotnet dotnet-maui dotnet-maui-blazor hacktoberfest
Something wrong? Category · Trend · Risk
@voicybot Telegram bot main repository
- Category
- multimodal media
- Stars
- 906
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
bot speech-to-text telegam telegram-bot
Something wrong? Category · Trend · Risk
Uncensored local AI studio for Windows, Linux, and macOS. Zero-setup GUI for Image Generation, GGUF LLMs, Text to Speech & Speech to Text
- Category
- multimodal media
- Stars
- 835
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +69 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
image-generation speech-to-text text-generation text-to-speech
Something wrong? Category · Trend · Risk
Whisper-Flow is a framework designed to enable real-time transcription of audio content using OpenAI’s Whisper model. Rather than processing entire files after upload (“batch mode”), Whisper-Flow acce
- Category
- multimodal media
- Stars
- 822
- Readiness
- ready (93/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
pypi-package python speech-to-text transcription whisper
Something wrong? Category · Trend · Risk
一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervention required.
- Category
- multimodal media
- Stars
- 791
- Readiness
- needs review (69/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alignment ass-subtitles captions diarization ffmpeg forced-alignment
Something wrong? Category · Trend · Risk
A modular Swift SDK for audio processing with MLX on Apple Silicon
- Category
- multimodal media
- Stars
- 749
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
mlx mlx-audio mlx-audio-swift mlx-swift-audio speech-to-speech speech-to-text
Something wrong? Category · Trend · Risk
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
- Category
- multimodal media
- Stars
- 735
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-ai audio-processing automatic-speech-recognition foundational-models large-audio-models large-language-model-speech
Something wrong? Category · Trend · Risk
说点啥(BiBi Keyboard):一个基于 Kotlin 的 Android 平台的 LLM 与 ASR 语音输入法键盘应用 An LLM ASR voice input method keyboard application for the Android platform based on Kotlin
- Category
- multimodal media
- Stars
- 731
- Readiness
- ready (90/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android-app asr doubao elevenlabs keyboard kotlin
Something wrong? Category · Trend · Risk
Free, open-source, 100% offline voice dictation for Linux. Speak and type anywhere via whisper.cpp, Whisper & VOSK engines, GPU-accelerated, works on X11 + Wayland!
- Category
- multimodal media
- Stars
- 722
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility dictation gpu-acceleration linux offline-first privacy-first
Something wrong? Category · Trend · Risk
Transcribe and translate voice into LRC file using Whisper and LLMs (GPT, Claude, et,al). 使用whisper和LLM(GPT,Claude等)来转录、翻译你的音频为字幕文件。
- Category
- multimodal media
- Stars
- 672
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auto-subtitle faster-whisper lyrics lyrics-generator openai-api openlrc
Something wrong? Category · Trend · Risk
On-device streaming speech-to-text engine powered by deep learning
- Category
- multimodal media
- Stars
- 670
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr automatic-speech-recognition online-speech-recognition speech-recognition speech-to-text streaming-speech-to-text
Something wrong? Category · Trend · Risk
Speech Recognition for React Native Expo projects
- Category
- multimodal media
- Stars
- 666
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +8 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
expo react-native speech-recognition speech-to-text voice-recognition
Something wrong? Category · Trend · Risk
A native macOS menu bar dictation app using local speech-to-text with WhisperKit
- Category
- multimodal media
- Stars
- 584
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon dictation macos menu-bar open-source privacy
Something wrong? Category · Trend · Risk
A Python library for solving reCAPTCHA v2 and v3 with Playwright
- Category
- multimodal media
- Stars
- 562
- Readiness
- ready (91/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asyncio library playwright recaptcha solver speech-to-text
Something wrong? Category · Trend · Risk
A modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChat
- Category
- multimodal media
- Stars
- 558
- Readiness
- ready (83/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chatbox haptics heartrate media node-programing osc
Something wrong? Category · Trend · Risk
Fast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Free and open-source.
- Category
- multimodal media
- Stars
- 545
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +17 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
apple-silicon dictation local-first macos meeting-recording neural-engine
Something wrong? Category · Trend · Risk
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
- Category
- multimodal media
- Stars
- 519
- Readiness
- ready (95/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +23 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cohere-transcribe cohere-transcribe-03-2026 ggml parakeet speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
open source audio and video transcription software
- Category
- multimodal media
- Stars
- 509
- Readiness
- ready (72/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
collaborative speech-to-text transcription
Something wrong? Category · Trend · Risk
On-device speech-to-text engine powered by deep learning
- Category
- multimodal media
- Stars
- 482
- Readiness
- needs review (70/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · high confidence
Why: 12 lifetime contributors
Why it may be a gem: open issue backlog is stable or shrinking
Strongest signals: push recency, contributor breadth, issue load. Risks: None identified. Missing inputs: None.
asr automatic-speech-recognition on-device speech-recognition speech-to-text stt
Something wrong? Category · Trend · Risk
Official Python SDK for Deepgram.
- Category
- multimodal media
- Stars
- 456
- Readiness
- ready (81/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 8/100 · high confidence
Why: +3 stars in 7 days; 45 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, contributor breadth, issue load. Risks: no pull-request review responses in 30 days. Missing inputs: None.
asr automated-speech-recognition deepgram python speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization (Dockerfile, CI image build and test)
- Category
- multimodal media
- Stars
- 454
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- medium
- Maintainer health
- healthy
- Maintenance risk
- 18/100 · high confidence
Why: 6 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: issue load, documentation, license. Risks: recent commit cadence is 1/4.2 of its monthly baseline. Missing inputs: release recency.
asr docker-image dockerfile speech speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignment, speech translation, voice i
- Category
- multimodal media
- Stars
- 447
- Readiness
- needs review (60/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 8/100 · high confidence
Why: +2 stars in 7 days; 12 commits in 30 days
Why it may be a gem: open issue backlog is stable or shrinking
Strongest signals: push recency, issue load, documentation. Risks: no pull-request review responses in 30 days. Missing inputs: None.
command-line forced-alignment language-detection language-identification node-js source-separation
Something wrong? Category · Trend · Risk
Open-source AI voice typing for macOS, Windows, and Linux. Press a hotkey, speak naturally, get polished text in any app.
- Category
- multimodal media
- Stars
- 438
- Readiness
- ready (78/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- healthy
- Maintenance risk
- 8/100 · high confidence
Why: +27 stars in 7 days; 48 commits in 30 days
Why it may be a gem: healthy maintenance and project fundamentals; consistent human and community activity
Strongest signals: push recency, commit activity, issue load. Risks: no pull-request review responses in 30 days. Missing inputs: None.
ai ai-tools byok cross-platform desktop-app dictation
Something wrong? Category · Trend · Risk
VRCT(VRChat Chatbox Translator & Transcription)
- Category
- multimodal media
- Stars
- 417
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
osc speech-recognition speech-to-text vrchat
Something wrong? Category · Trend · Risk
小牛视频翻译 是一款支持本地视频翻译、字幕翻译和 YouTube 视频翻译下载的 AI 工具,集成自动语音识别与多语言翻译功能,助力创作者高效完成视频翻译,应用于视频本地化与视频出海场景。
- Category
- multimodal media
- Stars
- 392
- Readiness
- ready (76/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-subtitles asr multilingual speech-recognition speech-to-text subtitle-translation
Something wrong? Category · Trend · Risk
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
- Category
- multimodal media
- Stars
- 381
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr automatic-speech-recognition onnx parakeet speaker-diarization speaker-identification
Something wrong? Category · Trend · Risk
Your personal voice interface for any app. Speak naturally and your words appear wherever your cursor is, with fully customizable AI voice dictation. Open source alternative to Wispr Flow.
- Category
- multimodal media
- Stars
- 374
- Readiness
- ready (74/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility cross-platform macos pipecat python rust
Something wrong? Category · Trend · Risk
A lightweight Python package for Automatic Speech Recognition using ONNX models
- Category
- multimodal media
- Stars
- 358
- Readiness
- ready (88/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr lightweight onnx parakeet python speech-recognition
Something wrong? Category · Trend · Risk
🎭 AI Avatar / digital human platform — upload a photo, clone a voice, talk to any face in real time with lip-sync video. Open-source, self-hosted. Claude · Whisper · Chatterbox · MuseTalk.
- Category
- multimodal media
- Stars
- 355
- Readiness
- ready (92/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, fork interest. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-avatar avatar-ai chatterbox-tts claude-ai digital-human fastapi
Something wrong? Category · Trend · Risk
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
- Category
- multimodal media
- Stars
- 26,771
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- risky
- Maintenance risk
- 100/100 · high confidence
Why: 100+ lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: contributor breadth, issue load, documentation. Risks: repository is archived, no push in 414 days, latest release is 2057 days old, 151 open issues with no maintainer issue responses in 30 days, no pull-request review responses in 30 days, no maintainer response activity in 30 days. Missing inputs: None.
Capped lower bounds: lifetime contributors.
deep-learning deepspeech embedded machine-learning neural-networks offline
Something wrong? Category · Trend · Risk
Faster Whisper transcription with CTranslate2
- Category
- multimodal media
- Stars
- 24,804
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- high
- Maintainer health
- risky
- Maintenance risk
- 38/100 · high confidence
Why: +139 stars in 7 days; 46 lifetime contributors
Why it may be a gem: healthy maintenance and project fundamentals; open issue backlog is stable or shrinking
Strongest signals: contributor breadth, issue load, documentation. Risks: no push in 261 days, no pull-request review responses in 30 days. Missing inputs: None.
deep-learning inference openai quantization speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
kaldi-asr/kaldi is the official location of the Kaldi project.
- Category
- multimodal media
- Stars
- 15,450
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 319 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
c-plus-plus cuda kaldi shell speaker-id speaker-verification
Something wrong? Category · Trend · Risk
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
- Category
- multimodal media
- Stars
- 15,030
- Readiness
- needs review (70/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +11 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android asr deep-learning deep-neural-networks deepspeech google-speech-to-text
Something wrong? Category · Trend · Risk
A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.
- Category
- multimodal media
- Stars
- 10,042
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +19 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
python realtime speech-to-text
Something wrong? Category · Trend · Risk
A Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
- Category
- multimodal media
- Stars
- 8,380
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 20/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 119 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asrt chinese-speech-recognition cnn ctc keras python
Something wrong? Category · Trend · Risk
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
- Category
- multimodal media
- Stars
- 5,616
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 27/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 165 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr speaker-diarization speech speech-recognition speech-to-text whisper
Something wrong? Category · Trend · Risk
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
- Category
- multimodal media
- Stars
- 4,730
- Readiness
- high risk (43/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +15 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 197 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech speech-recognition speech-to-text stt
Something wrong? Category · Trend · Risk
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
- Category
- multimodal media
- Stars
- 4,683
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 856 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning jax speech-recognition speech-to-text whisper
Something wrong? Category · Trend · Risk
The python library for real-time communication
- Category
- multimodal media
- Stars
- 4,620
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 207 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence hacktoberfest hacktoberfest2025 llm python real-time
Something wrong? Category · Trend · Risk
OpenAI Whisper ASR Webservice API
- Category
- multimodal media
- Stars
- 3,315
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 257 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr automatic-speech-recognition docker openai-whisper speech speech-recognition
Something wrong? Category · Trend · Risk
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
- Category
- multimodal media
- Stars
- 3,145
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 445 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
large-language-models multimodal-large-language-models speech-interaction speech-language-model speech-to-speech speech-to-text
Something wrong? Category · Trend · Risk
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
- Category
- multimodal media
- Stars
- 3,141
- Readiness
- high risk (40/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 273 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr ctranslate2 diarization faster-whisper openai speaker-diarization
Something wrong? Category · Trend · Risk
Lingvo
- Category
- multimodal media
- Stars
- 2,862
- Readiness
- needs review (65/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr distributed gpu-computing language-model lm machine-translation
Something wrong? Category · Trend · Risk
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
- Category
- multimodal media
- Stars
- 2,602
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 879 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr automatic-speech-recognition deep-learning speech-recognition speech-recognition-api speech-recognizer
Something wrong? Category · Trend · Risk
🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAI
- Category
- multimodal media
- Stars
- 2,367
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai artificial-intelligence awesome awesome-list gpt openai
Something wrong? Category · Trend · Risk
开源免费的 Wispr Flow 替代方案 | 集成FunASR本地模型和可配置大语言模型的下一代中文桌面语音工作流
- Category
- multimodal media
- Stars
- 2,256
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 303 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-text-processing chinese-speech-recognition electron-app funasr local-processing open-source
Something wrong? Category · Trend · Risk
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
- Category
- multimodal media
- Stars
- 2,173
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 933 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning neural-network speech-recognition speech-to-text stt tensorflow
Something wrong? Category · Trend · Risk
Free, easy, portable audio engine for games
- Category
- multimodal media
- Stars
- 2,143
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, documentation. Risks: no push in 724 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio blitzmax c cpp engine flac
Something wrong? Category · Trend · Risk
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
- Category
- multimodal media
- Stars
- 2,089
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +20 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
aitranslate hallucination japanese llm modelscope qwen3
Something wrong? Category · Trend · Risk
Voice activity detector (VAD) for the browser with a simple API
- Category
- multimodal media
- Stars
- 2,034
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +9 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 189 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
onnxruntime silero-vad speech-to-text typescript voice-activity-detection web
Something wrong? Category · Trend · Risk
an editor for spoken-word audio with automatic transcription
- Category
- multimodal media
- Stars
- 1,881
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-editing speech-to-text transcription video-editing
Something wrong? Category · Trend · Risk
Gathers machine learning and Tensorflow deep learning models for NLP problems, 1.13 < Tensorflow < 2.0
- Category
- multimodal media
- Stars
- 1,781
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 2209 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
attention chatbot deep-learning dnc-seq2seq embedded language-detection
Something wrong? Category · Trend · Risk
Kalliope is a framework that will help you to create your own personal assistant.
- Category
- multimodal media
- Stars
- 1,768
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 1132 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
bot bot-creation home-automation jarvis linux personal-assistant
Something wrong? Category · Trend · Risk
收集户晨风的所有内容
- Category
- multimodal media
- Stars
- 1,637
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 29/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 171 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
archiving content-analysis datasets internet-culture livestream social-media
Something wrong? Category · Trend · Risk
OBS plugin for local speech recognition and captioning using AI
- Category
- multimodal media
- Stars
- 1,569
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +7 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai live-streaming livestream obs obs-studio obs-studio-plugin
Something wrong? Category · Trend · Risk
Toolkit for efficient experimentation with Speech Recognition, Text2Speech and NLP
- Category
- multimodal media
- Stars
- 1,558
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 1914 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning float16 language-model mixed-precision multi-gpu multi-node
Something wrong? Category · Trend · Risk
the open-source virtual assistant for Ubuntu based Linux distributions
- Category
- multimodal media
- Stars
- 1,407
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1355 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
artificial-intelligence chatbot kaldi linux machine-learning nlp
Something wrong? Category · Trend · Risk
Synchronized Translation for Videos. Video dubbing
- Category
- multimodal media
- Stars
- 1,403
- Readiness
- needs review (60/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 17/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 102 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio-processing automatic-dubbing diarization document-translator dubbing
Something wrong? Category · Trend · Risk
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
- Category
- multimodal media
- Stars
- 1,398
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 792 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech-emotion-recognition speech-processing speech-recognition speech-separation speech-synthesis speech-to-text
Something wrong? Category · Trend · Risk
Whisper command line client compatible with original OpenAI client based on CTranslate2.
- Category
- multimodal media
- Stars
- 1,336
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 29/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 174 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
openai- openai-whisper speech-recognition speech-to-text whisper
Something wrong? Category · Trend · Risk
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
- Category
- multimodal media
- Stars
- 1,280
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 404 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
all-in-one asr audio-processing machine-translation non-autoregressive seamless
Something wrong? Category · Trend · Risk
A voice control - voice commands - speech recognition and speech synthesis javascript library. Create your own siri,google now or cortana with Google Chrome within your website.
- Category
- multimodal media
- Stars
- 1,269
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 1291 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
recognition speech-recognition speech-synthesis speech-to-text voice-commands
Something wrong? Category · Trend · Risk
A voice chat app
- Category
- multimodal media
- Stars
- 1,213
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai language-model python serverless speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
A TensorFlow Implementation of DC-TTS: yet another text-to-speech model
- Category
- multimodal media
- Stars
- 1,156
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, license. Risks: no push in 1211 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech speech-to-text tts
Something wrong? Category · Trend · Risk
AI Vtuber for Streaming on Youtube/Twitch
- Category
- multimodal media
- Stars
- 1,109
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +4 stars in 7 days
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-vtuber ai-waifu deepl openai speech-recognition speech-synthesis
Something wrong? Category · Trend · Risk
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
- Category
- multimodal media
- Stars
- 1,089
- Readiness
- needs review (46/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 232 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-to-text composable-architecture ios openai speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
💬📝 A small dictation app using OpenAI's Whisper speech recognition model.
- Category
- multimodal media
- Stars
- 1,089
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 713 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
dictation faster-whisper openai openai-api openai-whisper speech-recognition
Something wrong? Category · Trend · Risk
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
- Category
- multimodal media
- Stars
- 960
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 674 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai speech-recognition speech-to-text websocket
Something wrong? Category · Trend · Risk
Botium Speech Processing
- Category
- multimodal media
- Stars
- 943
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
botium speech-to-text text-to-speech
Something wrong? Category · Trend · Risk
Evaluate your speech-to-text system with similarity measures such as word error rate (WER)
- Category
- multimodal media
- Stars
- 922
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 19/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 113 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automatic-speech-recognition evaluation-metrics python3 speech-to-text wer word-error-rate
Something wrong? Category · Trend · Risk
An asynchronized Python library to automate solving ReCAPTCHA v2 using audio
- Category
- multimodal media
- Stars
- 894
- Readiness
- needs review (49/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 1156 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asyncio python recaptcha speech-to-text
Something wrong? Category · Trend · Risk
A free, open source, privacy-first voice input app for macOS.
- Category
- multimodal media
- Stars
- 886
- Readiness
- needs review (45/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: +6 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
accessibility dictation local-first macos menu-bar-app open-source
Something wrong? Category · Trend · Risk
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
- Category
- multimodal media
- Stars
- 872
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 233 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr chinese conformer deep-learning deepspeech2 paddlepaddle
Something wrong? Category · Trend · Risk
💬Speech recognition for your React app
- Category
- multimodal media
- Stars
- 842
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 427 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
react speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Build real time speech2text web apps using OpenAI's Whisper https://openai.com/blog/whisper/
- Category
- multimodal media
- Stars
- 834
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 329 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
machine-learning openai speech-recognition speech-to-text whisper
Something wrong? Category · Trend · Risk
The official repository of the Eesen project
- Category
- multimodal media
- Stars
- 834
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 2633 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr ctc ctc-loss kaldi speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Open STT
- Category
- multimodal media
- Stars
- 826
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, documentation. Risks: repository is archived, no push in 1610 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr automatic-speech-recognition dataset russian speech-to-text stt
Something wrong? Category · Trend · Risk
Descriptive Deep Learning
- Category
- multimodal media
- Stars
- 824
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 914 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
deep-learning deep-learning-tutorial deep-neural-networks image-recognition machine-learning neural-network
Something wrong? Category · Trend · Risk
🎤 Lobe TTS - A high-quality & reliable TTS/STT library for Server and Browser
- Category
- multimodal media
- Stars
- 800
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 26/100 · low confidence
Why: +2 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 158 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
auzre bun edge lobehub microsoft-speech-api nodejs
Something wrong? Category · Trend · Risk
Stephanie is an open-source platform built specifically for voice-controlled applications as well as to automate daily tasks imitating much of an virtual assistant's work.
- Category
- multimodal media
- Stars
- 797
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 2756 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
intent-prediction personal-assistant python speech-recognition speech-to-text stephanie
Something wrong? Category · Trend · Risk
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
- Category
- multimodal media
- Stars
- 795
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- watch
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chatbox discord free heart-rate osc speech-recognition
Something wrong? Category · Trend · Risk
Project that allows one to use a microphone with OpenAI whisper.
- Category
- multimodal media
- Stars
- 788
- Readiness
- needs review (60/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 764 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
microphone speech-recognition speech-to-text whisper whisper-ai whisper-api
Something wrong? Category · Trend · Risk
Local voice chatbot for engaging conversations, powered by Ollama, Hugging Face Transformers, and Coqui TTS Toolkit
- Category
- multimodal media
- Stars
- 788
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 725 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai assistant-chat-bots chatbot chatbots cli-app command-line-tool
Something wrong? Category · Trend · Risk
🎤 The easiest way to transcribe audio in Swift
- Category
- multimodal media
- Stars
- 784
- Readiness
- needs review (56/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 806 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ios macos openai speech-recognition speech-to-text swift
Something wrong? Category · Trend · Risk
A 100% private AI voice transcription app that converts speech to text in 100+ languages. Built with Compose Multiplatform for Android & iOS using Whisper AI - no cloud uploads, all processing happens
- Category
- multimodal media
- Stars
- 766
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +5 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android audio-player compose-ios compose-multiplatform compose-multiplatform-sample ios
Something wrong? Category · Trend · Risk
基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。
- Category
- multimodal media
- Stars
- 762
- Readiness
- needs review (59/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 233 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr chinese deep-learning deepspeech deepspeech2 docker
Something wrong? Category · Trend · Risk
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
- Category
- multimodal media
- Stars
- 750
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 477 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr openai speech-recognition speech-to-text stt unity3d
Something wrong? Category · Trend · Risk
Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。
- Category
- multimodal media
- Stars
- 729
- Readiness
- ready (75/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr conformer deep-learning deepspeech pytorch speech
Something wrong? Category · Trend · Risk
Private and on-device speech recognition keyboard and service for Android.
- Category
- multimodal media
- Stars
- 729
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 343 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android jetpack-compose keyboard kotlin kotlin-android material-3
Something wrong? Category · Trend · Risk
Adapt Intent Parser
- Category
- multimodal media
- Stars
- 718
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 747 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
intent-parser intents open-source opensource speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Speech to Text and KB input captions for OBS, VRChat, Twitch chat and Discord
- Category
- multimodal media
- Stars
- 714
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 780 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
captions obs speech-recognition speech-to-text tauri text-to-speech
Something wrong? Category · Trend · Risk
语音api示例
- Category
- multimodal media
- Stars
- 709
- Readiness
- needs review (47/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: limited evidence; inspect maintenance signals before adopting
Strongest signals: issue load, fork interest, documentation. Risks: no push in 743 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
baidu rest-api speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
This repository is deprecated. All of its content and history has been moved to googleapis/google-cloud-node.
- Category
- multimodal media
- Stars
- 682
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: repository is archived, no push in 1121 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
machine-learning nodejs speech speech-to-text
Something wrong? Category · Trend · Risk
A CLI script to generate subtitle files (SRT/VTT/TXT) for any video using either DeepSpeech or Coqui
- Category
- multimodal media
- Stars
- 651
- Readiness
- high risk (0/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 100/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: repository is archived, no push in 957 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr autosub coqui-ai deepspeech ffmpeg mozilla-deepspeech
Something wrong? Category · Trend · Risk
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
- Category
- multimodal media
- Stars
- 638
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 766 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
alexa hotword-detection keyword-spotting node speech speech-recognition
Something wrong? Category · Trend · Risk
Creating a software for automatic monitoring in online proctoring
- Category
- multimodal media
- Stars
- 631
- Readiness
- needs review (62/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 605 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automation dlib eye-tracking face-detection face-spoofing hacktoberfest
Something wrong? Category · Trend · Risk
Real-time transcription using faster-whisper
- Category
- multimodal media
- Stars
- 614
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 745 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
faster-whisper openai speech-recognition speech-to-text voice-recognition whisper
Something wrong? Category · Trend · Risk
An open-source on-device voice IME (keyboard) for Android using the Vosk library.
- Category
- multimodal media
- Stars
- 578
- Readiness
- high risk (42/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 402 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
android input-method-editor keyboard speech-to-text speech-to-text-android vosk
Something wrong? Category · Trend · Risk
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
- Category
- multimodal media
- Stars
- 577
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 710 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr deep-learning speech-recognition speech-to-text tensorrt tensorrt-llm
Something wrong? Category · Trend · Risk
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
- Category
- multimodal media
- Stars
- 556
- Readiness
- needs review (64/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
chatbots conversation conversational-bots conversational-interface conversational-ui input
Something wrong? Category · Trend · Risk
The J.A.R.V.I.S. Speech API is designed to be simple and efficient, using the speech engines created by Google to provide functionality for parts of the API. Essentially, it is an API written in Java,
- Category
- multimodal media
- Stars
- 542
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 2655 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
api google jarvis java recognition speech
Something wrong? Category · Trend · Risk
한국어 음성인식 STT API 리스트. 각 성능 벤치마크.
- Category
- multimodal media
- Stars
- 540
- Readiness
- high risk (44/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 0/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
awesome korean speech-recognition speech-to-text speech-to-text-api whisper
Something wrong? Category · Trend · Risk
This is a list of features, scripts, blogs and resources for better using Kaldi ( http://kaldi-asr.org/ )
- Category
- multimodal media
- Stars
- 536
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1640 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
automatic-speech-recognition awesome-list kaldi kaldi-asr speech speech-recognition
Something wrong? Category · Trend · Risk
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
- Category
- multimodal media
- Stars
- 527
- Readiness
- needs review (57/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 243 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr kaldi speech-recognition speech-to-text stt typescript
Something wrong? Category · Trend · Risk
Phonetisaurus G2P
- Category
- multimodal media
- Stars
- 518
- Readiness
- needs review (61/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 797 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
g2p lexicon openfst pronunciation-dictionary speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
This tool uses AI to evaluate your pronunciation.
- Category
- multimodal media
- Stars
- 512
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +3 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 356 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai language-learning pronunciation speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS
- Category
- multimodal media
- Stars
- 510
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 29/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 176 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
cuda deep-learning llama llm privacy speech-recognition
Something wrong? Category · Trend · Risk
VOSK Speech Recognition Toolkit
- Category
- multimodal media
- Stars
- 501
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1486 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
lifelong-learning multilingual python semi-supervised-learning speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
CSS10: A Collection of Single Speaker Speech Datasets for 10 Languages
- Category
- multimodal media
- Stars
- 491
- Readiness
- needs review (51/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, license. Risks: no push in 2345 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
dataset speech speech-to-text
Something wrong? Category · Trend · Risk
⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计。还内置了包括TikTok、抖音等社交媒体平台的爬虫,可实现来自多个社交平台的无缝媒体处理,为媒体内容数据自动化处理提供了强大且可扩展的解决方案。
- Category
- multimodal media
- Stars
- 470
- Readiness
- needs review (55/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 416 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr crawler douyin-api fastapi faster-whisper openai-whisper
Something wrong? Category · Trend · Risk
Record audio from a user's microphone and display a cool visualization.
- Category
- multimodal media
- Stars
- 470
- Readiness
- needs review (48/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, fork interest, documentation. Risks: no push in 621 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-recorder audio-visualizer microphone mp3-audio reactjs record-audio
Something wrong? Category · Trend · Risk
HuggingSound: A toolkit for speech-related tasks based on Hugging Face's tools
- Category
- multimodal media
- Stars
- 468
- Readiness
- needs review (53/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 1053 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio automatic-speech-recognition speech speech-recognition speech-to-text
Something wrong? Category · Trend · Risk
The dataset of Speech Recognition
- Category
- multimodal media
- Stars
- 466
- Readiness
- needs review (58/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 215 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr audio automatic-speech-recognition dataset deep-learning deep-neural-networks
Something wrong? Category · Trend · Risk
Fast text based video editing, node Electron Os X desktop app, with Backbone front end.
- Category
- multimodal media
- Stars
- 455
- Readiness
- needs review (54/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 887 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
autoedit backbone backbonejs desktop dmg edl
Something wrong? Category · Trend · Risk
FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
- Category
- multimodal media
- Stars
- 446
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: +1 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 925 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
audio-generation audio-quantization codec encodec speech-synthesis speech-to-text
Something wrong? Category · Trend · Risk
🇺🇦 Speech Recognition & Synthesis for Ukrainian
- Category
- multimodal media
- Stars
- 439
- Readiness
- high risk (39/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation. Risks: no push in 329 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
speech speech-recognition speech-synthesis speech-to-text text-to-speech tts
Something wrong? Category · Trend · Risk
Open source inference code for Rev's model
- Category
- multimodal media
- Stars
- 436
- Readiness
- needs review (52/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- risky
- Maintenance risk
- 30/100 · low confidence
Why: High-signal multimodal media project
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: issue load, documentation, license. Risks: no push in 472 days. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
asr asr-model canary deeplearning diarization docker
Something wrong? Category · Trend · Risk
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for every output.
- Category
- multimodal media
- Stars
- 564
- Readiness
- ready (86/100 heuristic points; not a probability)
- Data confidence
- low
- Maintainer health
- healthy
- Maintenance risk
- 0/100 · low confidence
Why: +65 stars in 7 days
Why it may be a gem: healthy maintenance and project fundamentals
Strongest signals: push recency, issue load, documentation. Risks: None identified. Missing inputs: commit activity, contributor breadth, release recency, response activity, maintenance distribution.
ai-pipeline audio-generation b2-labs backblaze backblaze-b2 data-pipeline
Something wrong? Category · Trend · Risk