Skip to content
77

AI Game DevTools (AI-GDT)

Your AI Game Dev Hub. The ultimate resource hub for AI-powered game development tools. Discover cutting-edge LLMs, World Model, Agent, Code, Image, Texture, Shader, 3D Model, Animation, Video, Audio, Music, Singing Voice and Analytics. 🔥

1.4k stars131 forks794 entriesLast push Jul 21, 2026 (2 months ago)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Project List >LLM (LLM & Tool)

AgentGPT

🤖 Assemble, configure, and deploy autonomous AI Agents in your browser.

In 9 listsDetails

AICommand

ChatGPT integration with Unity Editor.

In 4 listsDetails

AIOS

LLM Agent Operating System.

In 2 lists

AI Scientist

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

In 5 listsDetails

Assistant CLI

A comfortable CLI tool to use ChatGPT service🔥

In 2 lists

Auferet

An AI game master for solo text adventures and tabletop-style RPGs, with persistent memory of your story and your own uploaded lore.

Auto-GPT

An experimental open-source attempt to make GPT-4 fully autonomous.

In 7 listsDetails

BabyAGI

This Python script is an example of an AI-powered task management system.

In 8 listsDetails

👶🤖🖥️ BabyAGI UI

BabyAGI UI is designed to make it easier to run and develop with babyagi in a web app, like a ChatGPT.

In 3 lists

baichuan-7B

A large-scale 7B pretraining language model developed by Baichuan.

Baichuan-13B

A 13B large language model developed by Baichuan Intelligent Technology.

Baichuan 2

A series of large language models developed by Baichuan Intelligent Technology.

In 4 listsDetails

Bisheng

Bisheng is an open LLM devops platform for next generation AI applications.

In 5 listsDetails

Character-LLM

A Trainable Agent for Role-Playing.

ChatDev

Communicative Agents for Software Development.

In 6 listsDetails

ChatGPT-API-unity

Binds ChatGPT chat completion API to pure C# on Unity.

In 2 lists

ChatGPTForUnity

ChatGPT for unity.

ChatRWKV

ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.

In 3 lists

ChatYuan

Large Language Model for Dialogue in Chinese and English.

In 2 lists

Chinese-LLaMA-Alpaca-3

(Chinese Llama-3 LLMs) developed from Meta Llama 3.

Chrome-GPT

An AutoGPT agent that controls Chrome on your desktop.

In 2 lists

CogVLM

CogVLM, a powerful open-source visual language foundation model.

CoreNet

A library for training deep neural networks.

In 2 lists

Cosmos

Cosmos is a world model development platform that consists of world foundation models, tokenizers and video processing pipeline to accelerate the development of Physical AI at Robotics & AV labs.

In 3 lists

DBRX

DBRX is a large language model trained by Databricks.

DCLM

DataComp for Language Models.

DeepSeek-R1

DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning.

In 2 lists

DeepSeek-V3

DeepSeek-V3 is a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.

In 4 listsDetails

DemoGPT

Auto Gen-AI App Generator with the Power of Llama 2

In 9 listsDetails

Design2Code

Automating Front-End Engineering

Devika

Devika is an Agentic AI Software Engineer.

In 5 listsDetails

Devon

An open-source pair programmer.

In 3 lists

Dora

Generating powerful websites, one prompt at a time.

Flowise

Drag & drop UI to build your customized LLM flow using LangchainJS.

In 12 listsDetails

Gemini

Gemini is built from the ground up for multimodality — reasoning seamlessly across text, images, video, audio, and code.

In 2 lists

Gemma

Gemma is a family of lightweight, state-of-the art open models built from research and technology used to create Google Gemini models.

gemma.cpp

lightweight, standalone C++ inference engine for Google's Gemma models.

In 3 lists

GLM-4

GLM-4-9B is the open-source version of the latest generation of pre-trained models in the GLM-4 series launched by Zhipu AI.

In 2 lists

GLM-4.5

GLM-4.5: An open-source large language model designed for intelligent agents by Z.ai.

GPT4All

A chatbot trained on a massive collection of clean assistant data including code, stories and dialogue.

In 11 listsDetails

GPT-4o

GPT-4o (“o” for “omni”) is a step towards much more natural human-computer interaction—it accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs.

gpt-oss

gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI.

In 4 listsDetails

GPTScript

Develop LLM Apps in Natural Language.

In 2 lists

Grok-1

The weights and architecture of our 314 billion parameter Mixture-of-Experts model, Grok-1.

HuggingChat

Making the community's best AI chat models available to everyone.

In 5 listsDetails

Hugging Face API Unity Integration

This Unity package provides an easy-to-use integration for the Hugging Face Inference API, allowing developers to access and use Hugging Face AI models within their Unity projects.

In 2 lists

Hunyuan-MT

The Hunyuan-MT comprises a translation model, Hunyuan-MT-7B, and an ensemble model, Hunyuan-MT-Chimera. The translation model is used to translate source text into the target language, while the ensemble model integrates multiple translation outputs to produce a higher-quality result.

ImageBind

ImageBind One Embedding Space to Bind Them All.

In 2 lists

Index-1.9B

A SOTA lightweight multilingual LLM.

InteractML-Unity

InteractML, an Interactive Machine Learning Visual Scripting framework for Unity3D.

InteractML-Unreal Engine

Bringing Machine Learning to Unreal Engine.

InternLM

InternLM has open-sourced a 7 billion parameter base model, a chat model tailored for practical scenarios and the training system.

In 2 lists

InternLM-XComposer

InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension.

Jan

Bring AI to your Desktop.

In 8 listsDetails

Janus

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

In 2 lists

Kimi K2

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters.

In 3 lists

Lamini

Lamini allows any engineering team to outperform general purpose LLMs through RLHF and fine- tuning on their own data.

In 3 lists

LaMini-LM

LaMini-LM is a collection of small-sized, efficient language models distilled from ChatGPT and trained on a large-scale dataset of 2.58M instructions.

LangChain

LangChain is a framework for developing applications powered by language models.

In 6 listsDetails

LangFlow

⛓️ LangFlow is a UI for LangChain, designed with react-flow to provide an effortless way to experiment and prototype flows.

In 4 listsDetails

LaVague

Automate automation with Large Action Model framework.

In 2 lists

Lemur

Open Foundation Models for Language Agents.

Lepton AI

A Pythonic framework to simplify AI service building.

In 2 lists

Lit-LLaMA

Implementation of the LLaMA language model based on nanoGPT. Supports flash attention, Int8 and GPTQ 4bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training.

In 4 listsDetails

llama2-webui

Run Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac).

In 2 lists

Llama 3

The official Meta Llama 3 GitHub site.

In 2 lists

Llama 3.1

Llama is an accessible, open large language model (LLM) designed for developers, researchers, and businesses to build, experiment, and responsibly scale their generative AI ideas.

In 2 lists

LLaSM

Large Language and Speech Model.

LLM Answer Engine

Build a Perplexity-Inspired Answer Engine Using Next.js, Groq, Mixtral, Langchain, OpenAI, Brave & Serper.

In 3 lists

llm.c

LLM training in simple, raw C/CUDA.

LLMUnity

Create characters in Unity with LLMs!

LLocalSearch

LLocalSearch is a completely locally running search engine using LLM Agents.

In 4 lists

LogicGamesSolver

A Python tool to solve logic games with AI, Deep Learning and Computer Vision.

LongCat-Flash

LongCat-Flash is a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (averaging∼27B) based on contextual demands,…

LongWriter

LongWriter: Unleashing 10,000+ Word Generation From Long Context LLMs.

Large World Model (LWM)

Large World Model (LWM) is a general-purpose large-context multimodal autoregressive model.

Lumina-T2X

Lumina-T2X is a unified framework for Text to Any Modality Generation.

MetaGPT

The Multi-Agent Framework

In 8 listsDetails

MiniCPM-2B

An end-side LLM outperforms Llama2-13B.

In 3 lists

MiniGPT-4

Enhancing Vision-language Understanding with Advanced Large Language Models.

In 5 listsDetails

MiniGPT-5

Interleaved Vision-and-Language Generation via Generative Vokens.

MiniMax-01

MiniMax-01: Scaling Foundation Models with Lightning Attention.

Mixtral 8x7B

A high quality Sparse Mixture-of-Experts.

Mistral 7B

The best 7B model to date, Apache 2.0.

Mistral Large

Mistral Large is a new cutting-edge text generation model. It reaches top-tier reasoning capabilities.

MLC LLM

Enable everyone to develop, optimize and deploy AI models natively on everyone's devices.

In 4 listsDetails

MobiLlama

Towards Accurate and Lightweight Fully Transparent GPT.

In 2 lists

MoE-LLaVA

Mixture of Experts for Large Vision-Language Models.

Moshi

Moshi is an experimental conversational AI.

Moshi

Moshi: a speech-text foundation model for real time dialogue.

In 2 lists

MOSS

An open-source tool-augmented conversational language model from Fudan University.

In 3 lists

mPLUG-Owl🦉

Modularization Empowers Large Language Models with Multimodality.

In 2 lists

Nemotron-4

A 15-billion-parameter large multilingual language model trained on 8 trillion text tokens.

NExT-GPT

Any-to-Any Multimodal Large Language Model.

In 2 lists

OLMo

Open Language Model

In 2 lists

OmniLMM

Large multi-modal models for strong performance and efficient deployment.

OneLLM

One Framework to Align All Modalities with Language.

Open-Assistant

OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.

In 4 listsDetails

Open Deep Research

An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models.

In 3 lists

OpenDevin

An autonomous AI software engineer.

In 3 lists

Orion-14B

Orion-14B is a family of models includes a 14B foundation LLM, and a series of models.

Panda

Overseas Chinese open source large language model, based on Llama-7B, -13B, -33B, -65B for continuous pre-training in the Chinese field.

Perplexica

An AI-powered search engine.

In 4 listsDetails

Pi

AI chatbot designed for personal assistance and emotional support.

In 2 lists

Qwen1.5

Qwen1.5 is the improved version of Qwen.

Qwen2

Qwen2 is the large language model series developed by Qwen team, Alibaba Cloud.

Qwen2.5-Coder

Qwen2.5-Coder is the code version of Qwen2.5, the large language model series developed by Qwen team, Alibaba Cloud.

In 2 lists

Qwen-7B

The official repo of Qwen-7B (通义千问-7B) chat & pretrained large language model proposed by Alibaba Cloud.

Qwen3

Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.

In 5 listsDetails

RepoAgent

RepoAgent is an Open-Source project driven by Large Language Models(LLMs) that aims to provide an intelligent way to document projects.

In 3 lists

s1

s1: Simple test-time scaling.

In 2 lists

Sanity AI Engine

Sanity AI Engine for the Unity Game Development Tool.

SearchGPT

🌳 Connecting ChatGPT with the Internet

In 2 lists

Seed-OSS

Seed-OSS is a series of open-source large language models developed by ByteDance's Seed Team, designed for powerful long-context, reasoning, agent and general capabilities, and versatile developer-friendly features.

ShareGPT4V

Improving Large Multi-Modal Models with Better Captions.

SkyThought

Sky-T1: Train your own O1 preview model within $450.

Skywork

Skywork series models are pre-trained on 3.2TB of high-quality multilingual (mainly Chinese and English) and code data.

StableLM

Stability AI Language Models.

In 5 listsDetails

Stanford Alpaca

An Instruction-following LLaMA Model.

In 5 listsDetails

Text generation web UI

A gradio web UI for running Large Language Models like LLaMA, llama.cpp, GPT-J, OPT, and GALACTICA.

In 7 listsDetails

TinyChatEngine

On-Device LLM Inference Library.

ToolBench

An open platform for training, serving, and evaluating large language model for tool learning.

In 3 lists

Unity ChatGPT

Unity ChatGPT Experiments.

In 2 lists

Unity OpenAI-API Integration

Integrate openai GPT-3 language model and ChatGPT API into a Unity project.

Unreal Engine 5 Llama LoRA

A proof-of-concept project that showcases the potential for using small, locally trainable LLMs to create next-generation documentation tools.

UnrealGPT

A collection of Unreal Engine 5 Editor Utility widgets powered by GPT3/4.

Video-LLaVA

Learning United Visual Representation by Alignment Before Projection.

WebGPT

Run GPT model on the browser with WebGPU.

In 3 lists

Web3-GPT

Deploy smart contracts with AI

In 3 lists

WordGPT

🤖 Bring the power of ChatGPT to Microsoft Word

In 2 lists

XAgent

An Autonomous LLM Agent for Complex Task Solving.

In 5 listsDetails

Yi

A series of large language models trained from scratch by developers.

In 2 lists

01 Project

The open-source language model computer.

SimpleOllamaUnity

Ollama integration for Unity Engine (works in runtime and editor)

AI-Writer

AI writes novels, generates fantasy and romance web articles, etc. Chinese pre-trained generative model.

Notebook.ai

Notebook.ai is a set of tools for writers, game designers, and roleplayers to create magnificent universes – and everything within them.

Novel

Notion-style WYSIWYG editor with AI-powered autocompletions.

In 3 lists

NovelAI

Driven by AI, painlessly construct unique stories, thrilling tales, seductive romances, or just fool around.

In 2 lists

Unity-MCP

Open-source MCP server connecting AI agents to the Unity Editor and runtime, with 100+ built-in tools.

In 5 listsDetails

Godot-MCP

Open-source MCP server connecting AI agents to the Godot Editor and runtime (Godot 4.x, C#).

In 2 lists

Unreal-MCP

Open-source MCP server connecting AI agents to Unreal Engine 5.7, editor and runtime (C++ plugin + .NET sidecar).

In 2 lists

GameDev-MCP-Server

Open-source, engine-agnostic MCP server shared by Unity-MCP, Godot-MCP, and Unreal-MCP.

In 3 lists

MCP-Plugin-dotnet

Open-source .NET library/SDK that turns any .NET application into an MCP server.

ReflectorNet

Open-source .NET reflection toolkit for AI-driven scenarios.

VLM (Visual)

Cambrian-1

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

CogVLM2

GPT4V-level open-source multi-modal model based on Llama3-8B.

CoTracker

It is Better to Track Together.

In 2 lists

dots.vlm1

dots.vlm1 is the first vision-language model in the dots model family. Built upon a 1.2 billion-parameter vision encoder and the DeepSeek V3 large language model (LLM), dots.vlm1 demonstrates strong multimodal understanding and reasoning capabilities.

EVF-SAM

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

FaceHi

It is Better to Track Together.

GLM-V

GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

InternLM-XComposer

InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension.

Kangaroo

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Kwai Keye-VL

Kwai Keye-VL is a cutting-edge multimodal large language model meticulously crafted by the Kwai Keye Team at Kuaishou.

LGVI

Towards Language-Driven Video Inpainting via Multimodal Large Language Models.

LLaVA++

Extending Visual Capabilities with LLaMA-3 and Phi-3.

LLaVA-OneVision

LLaVA-OneVision: Easy Visual Task Transfer.

LongVA

Long Context Transfer from Language to Vision.

Lumina-DiMOO

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding.

MaskViT

Masked Visual Pre-Training for Video Prediction.

MiniCPM-Llama3-V 2.5

A GPT-4V Level MLLM on Your Phone.

In 4 listsDetails

MiniCPM-V 4.0

MiniCPM-V 4.0: A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone.

MoE-LLaVA

Mixture of Experts for Large Vision-Language Models.

MotionLLM

Understanding Human Behaviors from Human Motions and Videos.

PLLaVA

Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

POINTS-Reader

POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion.

Qwen-VL

A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Sapiens

Sapiens: Foundation for Human Vision Models.

ShareGPT4V

Improving Large Multi-modal Models with Better Captions.

SOLO

SOLO: A Single Transformer for Scalable Vision-Language Modeling.

VideoAgent

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

Video-CCAM

Video-CCAM: Advancing Video-Language Understanding with Causal Cross-Attention Masks.

Video-LLaVA

Learning United Visual Representation by Alignment Before Projection.

VideoLLaMA 2

Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoLLaMA 3

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-MME

The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Vitron

A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing.

VILA

VILA: On Pre-training for Visual Language Models.

Game (World Model & Agent)

AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents.

In 7 listsDetails

Agent Group Chat

An Interactive Group Chat Simulacra For Better Eliciting Collective Emergent Behavior.

Agent K

An autoagentic AGI that is self-evolving and modular.

In 2 lists

Agent Laboratory

Agent Laboratory: Using LLM Agents as Research Assistants.

In 2 lists

AgentScope

Start building LLM-empowered multi-agent applications in an easier way.

In 4 listsDetails

AgentSims

An Open-Source Sandbox for Large Language Model Evaluation.

AI Town

AI Town is a virtual town where AI characters live, chat and socialize.

In 4 listsDetails

anime.gf

Local & Open Source Alternative to CharacterAI.

Astrocade

Create games with AI

Atomic Agents

The Atomic Agents framework is designed to be modular, extensible, and easy to use.

AutoAgents

A Framework for Automatic Agent Generation.

In 2 lists

AutoGen

Enable Next-Gen Large Language Model Applications.

In 14 listsDetails

AWorld

AWorld: The Agent Runtime for Self-Improvement.

In 2 lists

behaviac

Behaviac is a framework of the game AI development.

In 3 lists

Biomes

Biomes is an open source sandbox MMORPG built for the web using web technologies such as Next.js, Typescript, React and WebAssembly.

Buffer of Thoughts

Thought-Augmented Reasoning with Large Language Models.

Byzer-Agent

Easy, fast, and distributed agent framework for everyone.

Cat Town

A C(h)atGPT-powered simulation with cats.

CharacterGLM

Customizing Chinese Conversational AI Characters with Large Language Models.

ChatDev

Communicative Agents for Software Development.

In 6 listsDetails

CogAgent

CogAgent is an open-source visual language model improved based on CogVLM.

ComoRAG

ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative Reasoning.

Cradle

Towards General Computer Control.

crewAI

Framework for orchestrating role-playing, autonomous AI agents.

In 4 listsDetails

Datarus Jupyter Agent

The Datarus Jupyter Agent is a powerful multi-step reasoning system that executes complex analytical workflows with step-by-step reasoning, automatic error recovery, and comprehensive result synthesis.

Dify

Dify is an open-source LLM app building platform.

In 14 listsDetails

Digital Life Project

Autonomous 3D Characters with Social Intelligence.

everything-ai

Your fully proficient, AI-powered and local chatbot assistant🤖.

fabric

fabric is an open-source framework for augmenting humans using AI.

In 6 listsDetails

FastGPT

FastGPT is a knowledge-based platform built on the LLM.

In 4 listsDetails

fastRAG

Efficient Retrieval Augmentation and Generation Framework.

GameAISDK

Image-based game AI automation framework.

GameNGen

Diffusion Models Are Real-Time Game Engines.

GameGen-O

GameGen-O: Open-world Video Game Generation.

GenAgent

GenAgent: Build Collaborative AI Systems with Automated Workflow Generation - Case Studies on ComfyUI.

Generative Agents

Interactive Simulacra of Human Behavior.

In 2 lists

Genesis

Genesis: A Generative and Universal Physics Engine for Robotics and Beyond.

Genie

Generative Interactive Environments.

In 2 lists

Genie 3

Genie 3: A new frontier for world models. Genie 3 is a general purpose world model that can generate an unprecedented diversity of interactive environments.

gigax

Runtime, LLM-powered NPCs.

HippoRAG

Neurobiologically Inspired Long-Term Memory for Large Language Models.

In 2 lists

Hunyuan-GameCraft

Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition.

HunyuanWorld 1.0

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels.

HunyuanWorld-Voyager

HunyuanWorld-Voyager is a novel video diffusion framework that generates world-consistent 3D point-cloud sequences from a single image with user-defined camera path. Voyager can generate 3D-consistent scene videos for world exploration following custom camera trajectories.

HY-World 1.5

HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency.

Interactive LLM Powered NPCs

Interactive LLM Powered NPCs, is an open-source project that completely transforms your interaction with non-player characters (NPCs) in any game!

IoA

An open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.

In 2 lists

Jaaz

Jaaz - The world's first open-source multimodal creative assistant. AI design agent, local alternative for Lovart. Canva + Cursor. AI agent with ability to design, edit and generate images, posters, storyboards, etc.

In 2 lists

KwaiAgents

A generalized information-seeking agent system with Large Language Models (LLMs).

In 2 lists

LangChain

Get your LLM application from prototype to production.

In 20 listsDetails

LangFlow

⛓️ LangFlow is a UI for LangChain, designed with react-flow to provide an effortless way to experiment and prototype flows.

In 4 listsDetails

LangGraph Studio

LangGraph Studio offers a new way to develop LLM applications by providing a specialized agent IDE that enables visualization, interaction, and debugging of complex agentic applications.

In 2 lists

LARP

Language-Agent Role Play for open-world games.

LLama Agentic System

Agentic components of the Llama Stack APIs.

In 3 lists

LlamaIndex

LlamaIndex is a data framework for your LLM application.

In 14 listsDetails

Matrix-Game

Matrix-Game: Interactive World Foundation Model. Matrix-Game is a 17B-parameter interactive world foundation model for controllable game world generation.

MindSearch

🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).

In 3 lists

Mixture of Agents (MoA)

Mixture-of-Agents Enhances Large Language Model Capabilities.

In 2 lists

MMRole

MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents.

Moonlander.ai

Start building 3D games without any coding using generative AI.

MuG Diffusion

MuG Diffusion is a charting AI for rhythm games based on Stable Diffusion (one of the most powerful AIGC models) with a large modification to incorporate audio waves.

NVIDIA NeMo Agent Toolkit

NVIDIA NeMo Agent toolkit is a flexible, lightweight, and unifying library that allows you to easily connect existing enterprise agents to data sources and tools across any framework.

In 2 lists

Oasis

Oasis is an interactive world model developed by Decart and Etched. Based on diffusion transformers, Oasis takes in user keyboard input and generates gameplay in an autoregressive manner.

OmAgent

A multimodal agent framework for solving complex tasks.

In 3 lists

OpenAgents

An Open Platform for Language Agents in the Wild.

In 7 listsDetails

OpenGame

OpenGame: Open Agentic Coding for Games.

In 2 lists

Opus

An AI app that turns text into a video game.

Pipecat

Open Source framework for voice and multimodal conversational AI.

In 7 listsDetails

Qwen-Agent

Qwen-Agent is a framework for developing LLM applications based on the instruction following, tool usage, planning, and memory capabilities of Qwen.

In 4 listsDetails

Ragas

Ragas is a framework that helps you evaluate your Retrieval Augmented Generation (RAG) pipelines.

In 3 lists

RPBench-Auto

An automated pipeline for evaluating LLMs for role-playing.

Rosebud AI

Vibe coding platform for creating 3D games and interactive web apps with AI.

In 3 lists

SIMA

A generalist AI agent for 3D virtual environments.

StoryGames.ai

AI for Dreamers Make Games.

SWE-agent

Agent Computer Interfaces Enable Software Engineering Language Models.

In 6 listsDetails

TaskGen

A Task-based agentic framework building on StrictJSON outputs by LLM agents.

TEN Agent

TEN Agent is the world’s first real-time multimodal agent integrated with the OpenAI Realtime API, RTC, and features weather checks, web search, vision, and RAG capabilities.

In 2 lists

Translation Agent

Agentic translation using reflection workflow.

Twitter

Twitter Personality is a web application that analyzes your Twitter handle to create a personalized personality profile using Wordware AI Agent.

Unbounded

Unbounded: A Generative Infinite Game of Character Life Simulation.

Video2Game

Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video.

V-IRL

Grounding Virtual Intelligence in Real Life.

WebDesignAgent

An agent used for webdesign.

XAgent

An Autonomous LLM Agent for Complex Task Solving.

In 5 listsDetails

Code

AI Code Translator

Use AI to translate code from one language to another.

In 2 lists

aiXcoder-7B

aiXcoder-7B Code Large Language Model.

bloop

bloop is a fast code search engine written in Rust.

In 4 listsDetails

Chapyter

ChatGPT Code Interpreter in Jupyter Notebooks.

In 3 lists

CodeGeeX

An Open Multilingual Code Generation Model.

In 2 lists

CodeGeeX2

A More Powerful Multilingual Code Generation Model.

CodeGeeX4

CodeGeeX4: Open Multilingual Code Generation Model.

CodeGen

CodeGen is an open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex.

In 4 listsDetails

CodeGen2

CodeGen2 models for program synthesis.

Code Llama

Code Llama is a large language models for code based on Llama 2.

In 2 lists

CodeTF

One-stop Transformer Library for State-of-the-art Code LLM.

CodeT5

Open Code LLMs for Code Understanding and Generation.

In 2 lists

Code World Model (CWM)

Code World Model (CWM) is a 32-billion-parameter open-weights LLM, to advance research on code generation with world models.

Cursor

Write, edit, and chat about your code with GPT-4 in a new type of editor.

In 3 lists

DeepSeek Coder

DeepSeek Coder: Let the Code Write Itself.

In 3 lists

OpenAI Codex

OpenAI Codex is a descendant of GPT-3.

In 2 lists

PandasAI

Pandas AI is a Python library that integrates generative artificial intelligence capabilities into Pandas, making dataframes conversational.

In 4 listsDetails

RobloxScripterAI

RobloxScripterAI is an AI-powered code generation tool for Roblox.

Roblox GUI Maker

Roblox GUI Maker generates Roblox Studio GUI layouts and Lua starter code from prompts for faster game UI prototyping.

Scikit-LLM

Seamlessly integrate powerful language models like ChatGPT into scikit-learn for enhanced text analysis tasks.

In 3 lists

SoTaNa

The Open-Source Software Development Assistant.

Stable Code 3B

Coding on the Edge.

StarCoder

💫 StarCoder is a language model (LM) trained on source code and natural language text.

StarCoder 2

StarCoder2 is a family of code generation models (3B, 7B, and 15B), trained on 600+ programming languages from The Stack v2 and some natural language text such as Wikipedia, Arxiv, and GitHub issues.

Tura

A terminal-native coding agent that turns intent into verified code changes with repo-aware controls and auditable execution.

In 7 listsDetails

UnityGen AI

UnityGen AI is an AI-powered code generation plugin for Unity.

Void

Void is an open source Cursor alternative. Write code with the best AI tools, retain full control over your data, and access powerful AI features.

In 2 lists

Image

AnyDoor

Zero-shot Object-level Image Customization.

AnyText

Multilingual Visual Text Generation And Editing.

AutoStudio

Crafting Consistent Subjects in Multi-turn Interactive Image Generation.

BAGEL

BAGEL - Unified Model for Multimodal Understanding and Generation. BAGEL is an open‑source multimodal foundation model with 7B active parameters (14B total) trained on large‑scale interleaved multimodal data.

In 2 lists

Blender-ControlNet

Using ControlNet right in Blender.

BriVL

Bridging Vision and Language Model.

CatVTON

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models.

CLIPasso

A method for converting an image of an object to a sketch, allowing for varying levels of abstraction.

ClipDrop

Create stunning visuals in seconds.

In 3 lists

ComfyUI

A powerful and modular stable diffusion GUI with a graph/nodes interface.

In 7 listsDetails

ConceptLab

Creative Generation using Diffusion Prior Constraints.

ControlNet

ControlNet is a neural network structure to control diffusion models by adding extra conditions.

In 3 lists

CSGO

CSGO: Content-Style Composition in Text-to-Image Generation.

DALL·E 2

DALL·E 2 is an AI system that can create realistic images and art from a description in natural language.

Dashtoon Studio

Dashtoon Studio is an AI powered comic creation platform.

DeepAI

DeepAI offers a suite of tools that use AI to enhance your creativity.

DeepFloyd IF

IF by DeepFloyd Lab at StabilityAI.

In 2 lists

Depth Anything V2

Depth Anything V2

Depth map library and poser

Depth map library for use with the Control Net extension for Automatic1111/stable-diffusion-webui.

Diffuse to Choose

Enriching Image Conditioned Inpainting in Latent Diffusion Models for Virtual Try-All.

Disco Diffusion

A frankensteinian amalgamation of notebooks, models and techniques for the generation of AI Art and Animations.

In 3 lists

DragGAN

Interactive Point-based Manipulation on the Generative Image Manifold.

In 4 listsDetails

Draw Things

AI- assisted image generation in Your Pocket.

DWPose

Effective Whole-body Pose Estimation with Two-stages Distillation.

EasyPhoto

Your Smart AI Photo Generator.

Flux

This repo contains minimal inference code to run text-to-image and image-to-image with our Flux latent rectified flow transformers.

In 3 lists

Follow-Your-Click

Open-domain Regional Image Animation via Short Prompts.

Fooocus

Focus on prompting and generating.

In 2 lists

GIFfusion

Create GIFs and Videos using Stable Diffusion.

Grounded-Segment-Anything

Automatically Detect , Segment and Generate Anything with Image, Text, and Audio Inputs.

HivisionIDPhotos

HivisionIDPhotos: a lightweight and efficient AI ID photos tools.

In 2 lists

Hua

Hua is an AI image editor with Stable Diffusion (and more).

Hunyuan-DiT

A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

In 2 lists

HunyuanImage-2.1

HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation​.

HunyuanImage-3.0

HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation​.

IC-Light

IC-Light is a project to manipulate the illumination of images.

Ideogram

Helping people become more creative.

Imagen

Imagen is an AI system that creates photorealistic images from input text.

In 4 listsDetails

img2img-turbo

One-Step Image-to-Image with SD-Turbo.

Img2Prompt

Get prompts from stable diffusion generated images.

Infinity

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

InstantID

Zero-shot Identity-Preserving Generation in Seconds.

InternLM-XComposer

InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension.

IRG

IRG - Interleaving Reasoning for Better Text-to-Image Generation.

KOALA

Self-Attention Matters in Knowledge Distillation of Latent Diffusion Models for Memory-Efficient and Fast Image Synthesis.

Kolors

Kolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis.

In 2 lists

Komiko

Komiko is an AI-powered storytelling platform that lets you create original characters, comics, and animations with ease.

KREA

Generate images and videos with a delightful AI-powered design tool.

In 4 listsDetails

LaVi-Bridge

Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation.

LayerDiffusion

Transparent Image Layer Diffusion using Latent Transparency.

Lexica

A Stable Diffusion prompts search engine.

In 5 listsDetails

LlamaGen

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Lumina-Image 2.0

Lumina-Image 2.0 : A Unified and Efficient Image Generative Model.

Lumina-mGPT

Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

MakeAnything

MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation.

MetaShoot

MetaShoot is a digital twin of a photo studio, developed as a plugin for Unreal Engine that gives any creator the ability to produce highly realistic renders in the easiest and quickest way.

Midjourney

Midjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species.

In 5 listsDetails

MIGC

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis.

MimicBrush

Zero-shot Image Editing with Reference Imitation.

NextStep-1

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale.

OmniGen

OmniGen: Unified Image Generation.

OmniGen2

OmniGen2: Exploration to Advanced Multimodal Generation.

In 2 lists

Oniichan

AI sprite generator and game character creator. Generate game-ready character sprites and original characters from text prompts using a custom finetuned model, with editing, inpainting, and a reusable character library.

Omost

Omost is a project to convert LLM's coding capability to image generation (or more accurately, image composing) capability.

Openpose Editor

Openpose Editor for AUTOMATIC1111's stable-diffusion-webui.

Outfit Anyone

Ultra-high quality virtual try-on for Any Clothing and Any Person.

PaintsUndo

PaintsUndo: A Base Model of Drawing Behaviors in Digital Paintings.

In 2 lists

PhotoMaker

Customizing Realistic Human Photos via Stacked ID Embedding.

Photoroom

AI Background Generator.

Plask

AI image generation in the cloud.

In 2 lists

PosterCraft

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework.

Prompt.Art

The Generators Hub.

PromptEnhancer

PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting.

PuLID

Pure and Lightning ID Customization via Contrastive Alignment.

In 2 lists

Qwen-Image

Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.

Rich-Text-to-Image

Expressive Text-to-Image Generation with Rich Text.

RPG-DiffusionMaster

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (PRG).

SEED-Story

SEED-Story: Multimodal Long Story Generation with Large Language Model.

Seedream AI Studio

Multi-model AI image generation using ByteDance Seedream 5.0/4.5/4.0 models, with one-click image-to-video animation via Kling 2.1. Free tier available.

In 2 lists

Segment Anything

Segment Anything Model (SAM): a new AI model from Meta AI that can "cut out" any object , in any image , with a single click.

In 2 lists

Segment Anything Model 2 (SAM 2)

SAM 2: Segment Anything in Images and Videos.

sd-webui-controlnet

WebUI extension for ControlNet.

In 2 lists

SDXL-Lightning

Progressive Adversarial Diffusion Distillation.

SDXS

Real-Time One-Step Latent Diffusion Models with Image Conditions.

SkyworkUniPic

SkyworkUniPic - Unified Autoregressive Modeling for Visual Understanding and Generation.

Sprite Fusion AI Pixel Art Generator

Generate, animate, edit, and export game-ready pixel art assets from text prompts.

In 2 lists

Sprite Fusion Pixel Snapper

Free and open-source tool that cleans AI-generated pixel art by restoring a consistent pixel grid and quantized color palette.

In 3 lists

Stable.art

Photoshop plugin for Stable Diffusion with Automatic1111 as backend (locally or with Google Colab).

Stable Cascade

Stable Cascade consists of three models: Stage A, Stage B and Stage C, representing a cascade for generating images, hence the name "Stable Cascade".

Stable Diffusion

A latent text-to-image diffusion model.

In 2 lists

stable-diffusion.cpp

Stable Diffusion in pure C/C++.

In 3 lists

Stable Diffusion web UI

A browser interface based on Gradio library for Stable Diffusion.

In 5 listsDetails

Stable Diffusion web UI

Web-based UI for Stable Diffusion.

Stable Diffusion WebUI Chinese

Chinese version of stable-diffusion-webui.

Stable Diffusion XL

Generate images from text.

Stable Diffusion XL Turbo

Real-Time Text-to-Image Generation.

Stable Diffusion 3.5

Stable Diffusion 3.5 open release includes multiple model variants, including Stable Diffusion 3.5 Large and Stable Diffusion 3.5 Large Turbo.

Stable Doodle

Stable Doodle is a sketch-to-image tool that converts a simple drawing into a dynamic image.

StableStudio

StableStudio by Stability AI

StoryMaker

StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation.

StreamDiffusion

A Pipeline-Level Solution for Real-Time Interactive Generation.

In 2 lists

StyleDrop

Text-To-Image Generation in Any Style.

SyncDreamer

Generating Multiview-consistent Images from a Single-view Image.

UltraEdit

UltraEdit: Instruction-based Fine-Grained Image Editing at Scale.

UltraPixel

UltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks.

Unity ML Stable Diffusion

Core ML Stable Diffusion on Unity.

USO

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning.

Vispunk Visions

Text-to-Image generation platform.

Texture

CRM

Single Image to 3D Textured Mesh with Convolutional Reconstruction Model.

DreamMat

High-quality PBR Material Generation with Geometry- and Light-aware Diffusion Models.

DreamSpace

Dreaming Your Room Space with Text-Driven Panoramic Texture Propagation.

Dream Textures

Stable Diffusion built-in to Blender. Create textures, concept art, background assets, and more with a simple text prompt.

In 4 lists

InstructHumans

Editing Animated 3D Human Textures with Instructions.

InteX

Interactive Text-to-Texture Synthesis via Unified Depth-aware Inpainting.

LLaMA-Mesh

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models.

MaterialSeg3D

MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets.

MeshAnything

MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D Assets.

Neuralangelo

High-Fidelity Neural Surface Reconstruction.

Paint-it

Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering.

Polycam

Create your own 3D textures just by typing.

TexFusion

Synthesizing 3D Textures with Text-Guided Image Diffusion Models.

Text2Tex

Text-driven texture Synthesis via Diffusion Models.

Texture Lab

AI-generated texures. You can generate your own with a text prompt.

With Poly

Create Textures With Poly. Generate 3D materials with AI in a free online editor, or search our growing community library.

X-Mesh

X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance.

Shader

AI Shader

ChatGPT-powered shader generator for Unity.

In 3 lists

3D Model

Animate3D

Animate3D: Animating Any 3D Model with Multi-view Video Diffusion.

Anything-3D

Segment-Anything + 3D. Let's lift the anything to 3D.

Any2Point

Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding.

BlenderGPT

Use commands in English to control Blender with OpenAI's GPT-4.

In 2 lists

Blender-GPT

An all-in-one Blender assistant powered by GPT3/4 + Whisper integration.

In 2 lists

BlenderMCP

BlenderMCP connects Blender to Claude AI through the Model Context Protocol (MCP), allowing Claude to directly interact with and control Blender. This integration enables prompt assisted 3D modeling, scene creation, and manipulation.

In 2 lists

Blockade Labs

Digital alchemy is real with Skybox Lab - the ultimate AI-powered solution for generating incredible 360° skybox experiences from text prompts.

In 2 lists

CF-3DGS

COLMAP-Free 3D Gaussian Splatting.

CharacterGen

CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.

chatGPT-maya

Simple Maya tool that utilizes open AI to perform basic tasks based on descriptive instructions.

CityDreamer

Compositional Generative Model of Unbounded 3D Cities.

CSM

Generate 3D worlds from images and videos.

Dash

Your Copilot for World Building in Unreal Engine.

Direct3D-S2

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention.

DreamCatalyst

DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation.

DreamGaussian4D

Generative 4D Gaussian Splatting.

DUSt3R

Geometric 3D Vision Made Easy.

Edify 3D

Edify 3D: Scalable High-Quality 3D Asset Generation.

GALA3D

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.

GaussCtrl

GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing.

GaussianCube

A Structured and Explicit Radiance Representation for 3D Generative Modeling.

GaussianDreamer

Fast Generation from Text to 3D Gaussian Splatting with Point Cloud Priors.

GenieLabs

Empower your game with AI-UGC.

HiFA

High-fidelity Text-to-3D with advance Diffusion guidance.

HoloDreamer

HoloDreamer: Holistic 3D Panoramic World Generation from Text Descriptions.

Hunyuan3D-1.0

Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation.

Hunyuan3D 2.0

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation.

Hunyuan3D 2.1

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material.

Infinigen

Infinite Photorealistic Worlds using Procedural Generation.

Instruct-NeRF2NeRF

Editing 3D Scenes with Instructions.

Interactive3D

Create What You Want by Interactive 3D Generation.

Isotropic3D

Image-to-3D Generation Based on a Single CLIP Embedding.

LATTE3D

Large-scale Amortized Text-To-Enhanced3D Synthesis.

LION

Latent Point Diffusion Models for 3D Shape Generation.

Luma AI

Capture in lifelike 3D. Unmatched photorealism, reflections, and details. The future of VFX is now, for everyone!

In 3 lists

lumine AI

AI-Powered Creativity.

Make-It-3D

High-Fidelity 3D Creation from A Single Image with Diffusion Prior.

Meshy

Create Stunning 3D Game Assets with AI.

In 3 lists

Mootion

Magical 3D AI Animation Maker.

MVDream

Multi-view Diffusion for 3D Generation.

NVIDIA Instant NeRF

Instant neural graphics primitives: lightning fast NeRF and more.

One-2-3-45

Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization.

Paint3D

Paint Anything 3D with Lighting-Less Texture Diffusion Models.

PAniC-3D

Stylized Single-view 3D Reconstruction from Portraits of Anime Characters.

PhysRig

PhysRig: Differentiable Physics-Based Rigging for Realistic Articulated Object Modeling.

Point·E

Point cloud diffusion for 3D model synthesis.

In 5 listsDetails

ProlificDreamer

High-Fidelity and diverse Text-to-3D generation with Variational score Distillation.

Seele AI

Input text to Generate playable 3D Games.

SF3D

SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement.

Shap-E

Generate 3D objects conditioned on text or images.

In 2 lists

Sloyd

3D modelling has never been easier.

In 2 lists

Spline AI

The power of AI is coming to the 3rd dimension. Generate objects, animations, and textures using prompts.

Stable Dreamfusion

A pytorch implementation of the text-to-3D model Dreamfusion, powered by the Stable Diffusion text-to-2D model.

In 4 lists

Step1X-3D

Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets.

SV3D

Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion.

Tafi

AI text to 3D character engine.

3D-GPT

Procedural 3D Modeling with Large Language Models.

3D-LLM

Injecting the 3D World into Large Language Models.

3Dpresso

Extract a 3D model of an object, captured on a video.

3DTopia

Text-to-3D Generation within 5 Minutes.

In 2 lists

3DTopia-XL

3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion.

threestudio

A unified framework for 3D content generation.

In 3 lists

TripoSR

A state-of-the-art open-source model for fast feedforward 3D reconstruction from a single image.

In 2 lists

Unique3D

High-Quality and Efficient 3D Mesh Generation from a Single Image.

In 2 lists

UnityGaussianSplatting

Toy Gaussian Splatting visualization in Unity.

ViVid-1-to-3

Novel View Synthesis with Video Diffusion Models.

Voxcraft

Crafting Ready-to-Use 3D Models with AI.

Wonder3D

Single Image to 3D using Cross-Domain Diffusion.

In 2 lists

Zero-1-to-3

Zero-shot One Image to 3D Object.

Avatar

AniPortrait

Audio-Driven Synthesis of Photorealistic Portrait Animations.

CALM

Conditional Adversarial Latent Models for Directable Virtual Characters.

ChatAvatar

Progressive generation Of Animatable 3D Faces Under Text guidance.

ChatdollKit

ChatdollKit enables you to make your 3D model into a chatbot.

In 5 listsDetails

Ditto

Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis.

DreamTalk

When Expressive Talking Head Generation Meets Diffusion Probabilistic Models.

Duix

Duix - Silicon-Based Digital Human SDK 🌐🤖

EchoMimic

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

EMOPortraits

Emotion-enhanced Multimodal One-shot Head Avatars.

EmoVOCA

EmoVOCA: Speech-Driven Emotional 3D Talking Heads.

E3 Gen

Efficient, Expressive and Editable Avatars Generation.

ExAvatar

ExAvatar - Expressive Whole-Body 3D Gaussian Avatar.

GeneAvatar

Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image.

GeneFace++

Generalized and Stable Real-Time 3D Talking Face Generation.

Hallo

Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

Hallo2

Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

HeadSculpt

Crafting 3D Head Avatars with Text.

HunyuanPortrait

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation.

HunyuanVideo-Avatar

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

IntrinsicAvatar

IntrinsicAvatar: Physically Based Inverse Rendering of Dynamic Humans from Monocular Videos via Explicit Ray Tracing.

Linly-Talker

Digital Avatar Conversational System.

LivePortrait

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

MotionGPT

Human Motion as a Foreign Language, a unified motion-language generation model using LLMs.

In 2 lists

MusePose

MusePose: a Pose-Driven Image-to-Video Framework for Virtual Human Generation.

MuseTalk

Real-Time High Quality Lip Synchorization with Latent Space Inpainting.

MuseV

Infinite-length and High Fidelity Virtual Human Video Generation with Visual Conditioned Parallel Denoising.

Portrait4D

Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data.

Ready Player Me

Integrate customizable avatars into your game or app in days.

In 2 lists

RodinHD

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models.

StableAvatar

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation.

StyleAvatar3D

Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation.

Text2Control3D

Controllable 3D Avatar Generation in Neural Radiance Fields using Geometry-Guided Text-to-Image Diffusion Model.

Topo4D

Topology-Preserving Gaussian Splatting for High-Fidelity 4D Head Capture.

UnityAIWithChatGPT

Based on Unity, ChatGPT+UnityChan voice interactive display is realized.

Vid2Avatar

3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene Decomposition.

VLOGGER

Multimodal Diffusion for Embodied Avatar Synthesis.

Wild2Avatar

Rendering Humans Behind Occlusions.

Animation

Animate Anyone

Consistent and Controllable Image-to-Video Synthesis for Character Animation.

AnimateAnything

Fine-Grained Open Domain Image Animation with Motion Guidance.

AnimateDiff

Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

In 4 lists

AnimateLCM

Let's Accelerate the Video Generation within 4 Steps!

Animate-X

Animate-X: Universal Character Image Animation with Enhanced Motion Representation.

AnimateZero

Video Diffusion Models are Zero-Shot Image Animators.

AnimationGPT

An AIGC tool for generating game combat motion assets.

Deforum

Deforum leverages Stable Diffusion to generate evolving AI visuals.

DrawingSpinUp

DrawingSpinUp: 3D Animation from Single Character Drawings.

DreaMoving

A Human Video Generation Framework based on Diffusion Models.

FaceFusion

Next generation face swapper and enhancer.

In 3 lists

FreeInit

Bridging Initialization Gap in Video Diffusion Models.

GeneFace

Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

ID-Animator

Zero-Shot Identity-Preserving Human Video Generation.

HY-Motion 1.0

HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation.

Index-AniSora

Index-AniSora is the most powerful open-source animated video generation model. It enables one-click creation of video shots across diverse anime styles including series episodes, Chinese original animations, manga adaptations, VTuber content, anime PVs, mad-style parodies(鬼畜动画), and more!

MagicAnimate

Temporally Consistent Human Image Animation using Diffusion Model.

NUWA

DragNUWA is an open-domain diffusion-based video generation model takes text, image, and trajectory controls as inputs to achieve controllable video generation.

NUWA-Infinity

NUWA-Infinity is a multimodal generative model that is designed to generate high-quality images and videos from given text, image or video input.

Omni Animation

AI Generated High Fidelity Animations.

PIA

Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models.

SadTalker

Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation.

SadTalker-Video-Lip-Sync

This project is based on SadTalkers Wav2lip for video lip synthesis.

Stable Animation

A powerful text-to-animation tool for developers.

ToonComposer

ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing.

TaleCrafter

An interactive story visualization tool that support multiple characters.

ToonCrafter

ToonCrafter: Generative Cartoon Interpolation.

In 2 lists

Wav2Lip

Accurately Lip-syncing Videos In The Wild.

In 3 lists

Wonder Studio

An AI tool that automatically animates, lights and composes CG characters into a live-action scene.

In 2 lists

Video

360DVD

Controllable Panorama Video Generation with 360-Degree Video Diffusion Model.

Animate-A-Story

Retrieval-Augmented Video Generation for Telling a Story.

Anything in Any Scene

Photorealistic Video Object Insertion.

ART•V

Auto-Regressive Text-to-Video Generation with Diffusion Models.

Assistive

Meet the generative video platform that brings your ideas to life.

AtomoVideo

High Fidelity Image-to-Video Generation.

BackgroundRemover

Background Remover lets you Remove Background from images and video using AI with a simple command line interface that is free and open source.

In 4 lists

Boximator

Generating Rich and Controllable Motions for Video Synthesis.

CoDeF

Content Deformation Fields for Temporally Consistent Video Processing.

CogVideo

Generate Videos from Text Descriptions.

CogVideoX

CogVideoX is an open-source version of the video generation model, which is homologous to 清影.

In 5 listsDetails

CogVLM

CogVLM is a powerful open-source visual language model (VLM).

CoNR

Genarate vivid dancing videos from hand-drawn anime character sheets(ACS).

Decohere

Create what can't be filmed.

Descript

Descript is the simple, powerful , and fun way to edit.

In 3 lists

Diffutoon

High-Resolution Editable Toon Shading via Diffusion Models.

In 3 lists

dolphin

General video interaction platform based on LLMs.

In 2 lists

DomoAI

Amplify Your Creativity with DomoAI.

DreamCinema

DreamCinema: Cinematic Transfer with Free Camera and 3D Character.

DynamiCrafter

Animating Open-domain Images with Video Diffusion Priors.

EDGE

We introduce EDGE, a powerful method for editable dance generation that is capable of creating realistic, physically-plausible dances while remaining faithful to arbitrary input music.

EMO

Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

Emu Video

Factorizing Text-to-Video Generation by Explicit Image Conditioning.

In 3 lists

Etna

Etna can generate corresponding video content based on short text descriptions.

Fairy

Fast Parallelized Instruction-Guided Video-to-Video Synthesis.

Follow-Your-Canvas

Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation.

Follow Your Pose

Pose-Guided Text-to-Video Generation using Pose-Free Videos.

FullJourney

Your complete suite of AI Creation tools at your fingertips.

Gen-2

A multi-modal AI system that can generate novel videos with text, images, or video clips.

In 2 lists

Generative Dynamics

Generative Image Dynamics.

Genie

Generative Interactive Environments.

In 2 lists

Genmo

Magically make videos with AI.

GenTron

Diffusion Transformers for Image and Video Generation.

HiGen

Hierarchical Spatio-temporal Decoupling for Text-to-Video generation.

Hotshot-XL

Hotshot-XL is an AI text-to-GIF model trained to work alongside Stable Diffusion XL.

In 2 lists

HuMo

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning.

HunyuanVideo

HunyuanVideo: A Systematic Framework For Large Video Generation Model.

In 4 listsDetails

HunyuanVideo-1.5

HunyuanVideo-1.5: A leading lightweight video generation model.

In 2 lists

Imagen Video

Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of interleaved spatial and temporal video super-resolution models.

InfiniteTalk

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing.

InstructVideo

Instructing Video Diffusion Models with Human Feedback.

I2VGen-XL

High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

LaVie

High-Quality Video Generation with Cascaded Latent Diffusion Models.

LongLive

LongLive: Real-time Interactive Long Video Generation.

In 2 lists

LTX Studio

LTX Studio is a holistic, AI-driven filmmaking platform for creators, marketers, filmmakers and studios.

In 2 lists

LTX-Video

LTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time. It can generate 24 FPS videos at 768x512 resolution, faster than it takes to watch them.

In 5 listsDetails

Lumiere

A Space-Time Diffusion Model for Video Generation.

LVDM

Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Lynx

Lynx: Towards High-Fidelity Personalized Video Generation.

MagicVideo

Efficient Video Generation With Latent Diffusion Models.

MagicVideo-V2

Multi-Stage High-Aesthetic Video Generation.

Magic Hour

AI Video for Creators made simple.

MAGVIT-v2

Tokenizer is key to visual generation.

MAGVIT

Masked Generative Video Transformer.

Make-A-Video

Make-A-Video is a state-of-the-art AI system that generates videos from text.

Make Pixels Dance

High-Dynamic Video Generation.

Make-Your-Video

Customized Video Generation Using Textual and Structural Guidance.

MicroCinema

A Divide-and-Conquer Approach for Text-to-Video Generation.

MIMO

MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling.

Mini-Gemini

Mining the Potential of Multi-modality Vision Language Models.

MobileVidFactory

Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text.

In 2 lists

Mochi 1

Mochi 1 is an open state-of-the-art video generation model with high-fidelity motion and strong prompt adherence in preliminary evaluation.

MOFA-Video

Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model.

MoneyPrinterTurbo

Use large models to generate short videos with one click.

In 4 listsDetails

Moonvalley

Moonvalley is a groundbreaking new text-to-video generative AI model.

Mora

More like Sora for Generalist Video Generation.

Morph Studio

With our Text-to-Video AI Magic, manifest your creativity through your prompt.

MotionClone

MotionClone: Training-Free Motion Cloning for Controllable Video Generation.

MotionCtrl

A Unified and Flexible Motion Controller for Video Generation.

MotionDirector

Motion Customization of Text-to-Video Diffusion Models.

Motionshop

An application of replacing the characters in video with 3D avatars.

Mov2mov

Mov2mov plugin for Automatic1111/stable-diffusion-webui.

MovieFactory

Automatic Movie Creation from Text using Large Generative Models for Language and Images.

MoviiGen 1.1

MoviiGen 1.1: Towards Cinematic-Quality Video Generative Models. MoviiGen 1.1 is a cutting-edge video generation model that excels in cinematic aesthetics and visual quality. This model is a fine-tuning model based on the Wan2.1. Based on comprehensive evaluations by 11 professional filmmakers and…

Neural Frames

Discover the synthesizer for the visual world.

NeverEnds

Create your world.

In 2 lists

Open-Sora

Democratizing Efficient Video Production for All.

In 4 listsDetails

Open-Sora

Open-Sora Plan.

In 4 listsDetails

Ovi

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Phenaki

A model for generating videos from text, with prompts that can change over time, and videos that can be as long as multiple minutes.

Pika Labs

Pika Labs is revolutionizing video-making experience with AI.

In 5 listsDetails

Pixeling

Pixeling empowers our customers to create highly precise, ultra-realistic, and extremely controllable visual content including images, videos and 3D models.

PixVerse

Create breath-taking videos with AI.

Pollinations

Creating gets easy, fast, and fun.

Reuse and Diffuse

Iterative Denoising for Text-to-Video Generation.

Ruyi

Ruyi is an image-to-video model capable of generating cinematic-quality videos at a resolution of 768, with a frame rate of 24 frames per second, totaling 5 seconds and 120 frames.

ShortGPT

An experimental AI framework for automated short/video content creation.

In 5 listsDetails

Show-1

Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

Step-Video-T2V

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

In 2 lists

SkyReels-A1

SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers.

SkyReels-V1

SkyReels V1: Human-Centric Video Foundation Model.

Snap Video

Scaled Spatiotemporal Transformers for Text-to-Video Synthesis.

Sora

Creating video from text.

In 3 lists

SoraWebui

SoraWebui is an open-source Sora web client, enabling users to easily create videos from text with OpenAI's Sora model.

StableVideo

Text-driven Consistency-aware Diffusion Video Editing.

Stable Video Diffusion

Stable Video Diffusion (SVD) Image-to-Video.

In 2 lists

StoryDiffusion

Consistent Self-Attention for Long-Range Image and Video Generation.

StoryMem

StoryMem: Multi-shot Long Video Storytelling with Memory.

StreamingT2V

Consistent, Dynamic, and Extendable Long Video Generation from Text.

StyleCrafter

nhancing Stylized Text-to-Video Generation with Style Adapter.

TATS

Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer.

Text2Video-Zero

Text-to-Image Diffusion Models are Zero-Shot Video Generators.

In 3 lists

TF-T2V

A Recipe for Scaling up Text-to-Video Generation with Text-free Videos.

Tora

Tora: Trajectory-oriented Diffusion Transformer for Video Generation.

Track-Anything

Track-Anything is a flexible and interactive tool for video object tracking and segmentation, based on Segment Anything and XMem.

In 2 lists

Tune-A-Video

One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation.

TwelveLabs

Multimodal AI that understands videos like humans.

In 2 lists

UniVG

Towards UNIfied-modal Video Generation.

Vchitect-2.0

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

VGen

A holistic video generation ecosystem for video generation building on diffusion models.

ViewCrafter

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.

Video-ChatGPT

Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos.

In 2 lists

VideoComposer

Compositional Video Synthesis with Motion Controllability.

VideoCrafter1

Open Diffusion Models for High-Quality Video Generation.

In 2 lists

VideoCrafter2

Overcoming Data Limitations for High-Quality Video Diffusion Models.

VideoDrafter

Content-Consistent Multi-Scene Video Generation with LLM.

VideoElevator

Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models.

VideoFactory

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

In 2 lists

VideoGen

A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

VideoLCM

Video Latent Consistency Model.

In 2 lists

Video LDMs

Align your Latents: High- resolution Video Synthesis with Latent Diffusion Models.

In 2 lists

Video-LLaVA

Learning United Visual Representation by Alignment Before Projection.

VideoMamba

State Space Model for Efficient Video Understanding.

Video-of-Thought

Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition.

VideoPoet

A large language model for zero-shot video generation.

Vispunk Motion

Create realistic videos using just text.

VisualRWKV

VisualRWKV is the visual-enhanced version of the RWKV language model, enabling RWKV to handle various visual tasks.

V-JEPA

Video Joint Embedding Predictive Architecture.

W.A.L.T

Photorealistic Video Generation with Diffusion Models.

Wan2.1

Wan: Open and Advanced Large-Scale Video Generative Models.

In 4 listsDetails

Wan2.2

Wan: Open and Advanced Large-Scale Video Generative Models.

In 3 lists

Waver

Waver 1.0 is a next-generation, universal foundation model family for unified image and video generation, built on rectified flow Transformers and engineered for industry-grade performance.

Zeroscope

Zeroscope Text-to-Video.

Audio

AcademiCodec

An Open Source Audio Codec Model for Academic Research.

Amphion

An Open-Source Audio, Music, and Speech Generation Toolkit.

In 2 lists

ArchiSound

Audio generation using diffusion models, in PyTorch.

In 2 lists

Audiobox

Unified Audio Generation with Natural Language Prompts.

AudioEditing

Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

Audiogen Codec

A low compression 48khz stereo neural audio codec for general audio, optimizing for audio fidelity 🎵.

AudioGPT

Understanding and Generating Speech, Music, Sound, and Talking Head.

In 4 listsDetails

AudioLCM

Text-to-Audio Generation with Latent Consistency Models.

AudioLDM

Text-to-Audio Generation with Latent Diffusion Models.

In 2 lists

AudioLDM 2

Learning Holistic Audio Generation with Self-supervised Pretraining.

AudioX

AudioX: Diffusion Transformer for Anything-to-Audio Generation.

Auffusion

Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

CTAG

Creative Text-to-Audio Generation via Synthesizer Programming.

FoleyCrafter

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

HunyuanVideo-Foley

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

MAGNeT

Masked Audio Generation using a Single Non-Autoregressive Transformer.

Make-An-Audio

Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.

Make-An-Audio 3

Transforming Text into Audio via Flow-based Large Diffusion Transformers.

MeanAudio

MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows.

MiDashengLM

MiDashengLM: Efficient Audio Understanding with General Audio Captions.

MMAudio

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

NeuralSound

Learning-based Modal Sound Synthesis with Acoustic Transfer.

OptimizerAI

Sounds for Creators, Game makers, Artists, Video makers.

Qwen2-Audio

Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.

SEE-2-SOUND

Zero-Shot Spatial Environment-to-Spatial Sound.

SoundStorm

Efficient Parallel Audio Generation.

Stable Audio

Fast Timing-Conditioned Latent Audio Diffusion.

In 3 lists

Stable Audio Open

Stable Audio Open 1.0 generates variable-length (up to 47s) stereo audio at 44.1kHz from text prompts.

SyncFusion

SyncFusion: Multimodal Onset-synchronized Video-to-Audio Foley Synthesis.

TANGO

Text-to-Audio Generation using Instruction Tuned LLM and Latent Diffusion Model.

ThinkSound

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing.

VTA-LDM

Video-to-Audio Generation with Hidden Alignment.

WavJourney

Compositional Audio Creation with Large Language Models.

Music

AIVA

The Artificial Intelligence composing emotional soundtrack music.

In 5 listsDetails

Amper Music

Custom music generation technology powered by Amper.

AnyAccomp

AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck.

Boomy

Create generative music. Share it with the world.

In 2 lists

ChatMusician

Fostering Intrinsic Musical Abilities Into LLM.

Chord2Melody

Automatic Music Generation AI.

Diff-BGM

A Diffusion Model for Video Background Music Generation.

FluxMusic

FluxMusic: Text-to-Music Generation with Rectified Flow Transformer.

GPTAbleton

Draft script for processing GPT response and sending the MIDI notes into the Ableton clips with AbletonOSC and python-osc.

HeyMusic.AI

AI Music Generator

Image to Music

AI Image to Music Generator is a tool that uses artificial intelligence to convert images into music.

JEN-1

Text-Guided Universal Music Generation with Omnidirectional Diffusion Models.

Jukebox

A Generative Model for Music.

In 2 lists

Magenta

Magenta is a research project exploring the role of machine learning in the process of creating art and music.

In 3 lists

MeLoDy

Efficient Neural Music Generation

Mubert

AI Generative Music.

In 3 lists

MuseNet

A deep neural network that can generate 4-minute musical compositions with 10 different instruments, and can combine styles from country to Mozart to the Beatles.

MusicGen

Simple and Controllable Music Generation.

In 3 lists

MusicLDM

Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.

MusicLM

Generating Music From Text.

In 4 listsDetails

Riffusion App

Riffusion is an app for real-time music generation with stable diffusion.

Sonauto

Sonauto is an AI music editor that turns prompts, lyrics, or melodies into full songs in any style.

SonicMaster

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering.

SoundRaw

AI music generator for creators.

In 3 lists

Soundry AI

Generative AI tools including text-to-sound and infinite sample packs.

YuE

YuE: Open Full-song Generation Foundation Model, something similar to Suno.ai but open.

In 3 lists

Singing Voice

DiffSinger

Singing Voice Synthesis via Shallow Diffusion Mechanism.

Retrieval-based-Voice-Conversion-WebUI

An easy-to-use SVC framework based on VITS.

so-vits-svc

SoftVC VITS Singing Voice Conversion.

In 2 lists

VI-SVS

Use VITS and Opencpop to develop singing voice synthesis; Different from VISinger.

Speech

Applio

Ultimate voice cloning tool, meticulously optimized for unrivaled power, modularity, and user-friendly experience.

Audyo

Text in. Audio out.

Bark

Text-Prompted Generative Audio Model.

In 6 listsDetails

Bert-VITS2

VITS2 Backbone with multilingual bert.

Chatterbox

Chatterbox TTS is the first production-grade open-source TTS model.

In 3 lists

ChatTTS

ChatTTS is a generative speech model for daily dialogue.

In 4 listsDetails

CLAPSpeech

Learning Prosody from Text Context with Contrastive Language-Audio Pre-Training.

CosyVoice

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

DEX-TTS

Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability.

EmotiVoice

A Multi-Voice and Prompt-Controlled TTS Engine.

In 2 lists

FireRedTTS-2

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

In 2 lists

Fliki

Turn text into videos with AI voices.

In 3 lists

GLM-4-Voice

GLM-4-Voice is an end-to-end voice model launched by Zhipu AI. GLM-4-Voice can directly understand and generate Chinese and English speech, engage in real-time voice conversations, and change attributes such as emotion, intonation, speech rate, and dialect based on user instructions.

Glow-TTS

A Generative Flow for Text-to-Speech via Monotonic Alignment Search.

GPT-SoVITS

A Powerful Few-shot Voice Conversion and Text-to-Speech WebUI.

In 3 lists

Higgs Audio

Higgs Audio V2: Redefining Expressiveness in Audio Generation.

In 2 lists

IndexTTS2

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech.

In 2 lists

Kitten TTS

Kitten TTS is an open-source realistic text-to-speech model with just 15 million parameters, designed for lightweight deployment and high-quality voice synthesis.

In 4 listsDetails

Liquid Audio

Liquid Audio - Speech-to-Speech audio models by Liquid AI.

LOVO

LOVO is the go-to AI Voice Generator & Text to Speech platform for thousands of creators.

In 3 lists

MahaTTS

An Open-Source Large Speech Generation Model.

Matcha-TTS

A fast TTS architecture with conditional flow matching.

MeloTTS

High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

In 2 lists

MetaVoice-1B

AI for human-level speech intelligence.

Narakeet

Easily Create Voiceovers Using Realistic Text to Speech.

Mini-Omni

Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming. Mini-Omni is an open-source multimodel large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.

One-Shot-Voice-Cloning

One Shot Voice Cloning base on Unet-TTS.

OpenVoice

Instant voice cloning by MyShell.

In 2 lists

OverFlow

Putting flows on top of neural transducers for better TTS.

RealtimeTTS

RealtimeTTS is a state-of-the-art text-to-speech (TTS) library designed for real-time applications.

SenseVoice

SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).

In 2 lists

SpeechGPT

Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

speech-to-text-gpt3-unity

This is the repo I use Whisper and ChatGPT API from OpenAI in Unity.

Stable Speech

Stability AI's Text-to-Speech model.

StableTTS

Next-generation TTS model using flow-matching and DiT, inspired by Stable Diffusion 3.

Step-Audio

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Step-Audio 2

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

In 2 lists

StyleTTS 2

Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.

tortoise.cpp

tortoise.cpp: GGML implementation of tortoise-tts.

TorToiSe-TTS

A multi-voice TTS system trained with an emphasis on quality.

In 5 listsDetails

TTS Generation WebUI

TTS Generation WebUI (Bark, MusicGen, Tortoise, RVC, Vocos, Demucs).

In 2 lists

VALL-E

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

In 2 lists

VALL-E X

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

In 2 lists

VibeVoice

VibeVoice is a novel framework designed for generating expressive, long-form, multi-speaker conversational audio, such as podcasts, from text. It addresses significant challenges in traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking.

In 5 listsDetails

Vocode

Vocode is an open-source library for building voice-based LLM applications.

Voicebox

Text-Guided Multilingual Universal Speech Generation at Scale.

VoiceCraft

Zero-Shot Speech Editing and Text-to-Speech in the Wild.

VoxCPM

VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning.

In 4 listsDetails

Whisper

Whisper is a general-purpose speech recognition model.

In 11 listsDetails

WhisperSpeech

An Open Source text-to-speech system built by inverting Whisper.

X-E-Speech

Joint Training Framework of Non-Autoregressive Cross-lingual Emotional Text-to-Speech and Voice Conversion.

XTTS

XTTS is a library for advanced Text-to-Speech generation.

In 4 listsDetails

YourTTS

Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone.

ZMM-TTS

Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations.

UniAudio 2.0

UniAudio 2.0: A Multi-task Audio Foundation Model with Reasoning-Augmented Audio Tokenization.

UnityNeuroSpeech

The world’s first game framework that lets you talk to AI in real time — locally.

Analytics

Ludo.ai

Assistant for game research and design.

In 2 lists
See category
94

Awesome Mac

jaywcjlove/awesome-mac

 This project is dedicated to collecting high-quality macOS software and organizing them systematically by different categories for easy search and use.

Fresh★ 115k1316 entriesPushed today
91

Open Source Mac Os Apps

serhii-londar/open-source-mac-os-apps

🚀 Awesome list of open source applications for macOS. https://t.me/s/opensourcemacosapps

Fresh★ 51k700 entriesPushed 20 days ago
91

Awesome-Kubernetes

ramitsurana/awesome-kubernetes

A curated list for awesome kubernetes sources :ship::tada:

Fresh★ 16k47 entriesPushed 8 days ago
90

Awesome Nodejs

sindresorhus/awesome-nodejs

:zap: Delightful Node.js packages and resources [BECAUSE OF TOO MUCH SPAM AND LOW-QUALITY SUBMISSIONS, SUBMISSIONS ARE PAUSED TEMPORARILY]

Fresh★ 67k588 entriesPushed 28 days ago
90

Awesome Home Assistant

frenck/awesome-home-assistant

A curated list of amazingly awesome Home Assistant resources.

Fresh★ 8.5k312 entriesPushed 2 days ago
90

Awesome Ios

vsouza/awesome-ios

A curated list of awesome iOS ecosystem, including Objective-C and Swift Projects

Fresh★ 53k1812 entriesPushed 1 month ago