When LLM Agents Meet Reinforcement Learning
Section: Base Framework · HuggingFace
Entry
Appears in 8 awesome lists
Train transformer language models with reinforcement learning.
Section: Base Framework · HuggingFace
Section: Foundation Model Fine Tuning · Train transformer language models with reinforcement learning.
Section: Training and Fine-tuning · train transformer language models with reinforcement learning
Section: Instruction Tuning and Preference Optimization · reference library for SFT, DPO, GRPO, and RLHF.
Section: 7. Training & Fine-tuning Ecosystem · Official library for RLHF, SFT, DPO, ORPO.
Section: Industry Strength Reinforcement Learning · Train transformer language models with reinforcement learning.
Section: Graph Machine Learning · Train transformer language models with reinforcement learning.
Section: Other · Train transformer language models with reinforcement learning.
Parameter-Efficient Fine-Tuning (PEFT) methods enable efficient adaptation of pre-trained language models (PLMs) to various downstream applications without fine-tuning all the model's parameters.
LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible
Extensible toolkit for finetuning and inference of large foundation models. Features RAFT alignment algorithm and comprehensive model support. Apache 2.0 licensed.
Code for rproducing the Stanford Alpaca results using low-rank adaptation (LoRA).
Scalable open-source RL infrastructure for post-training foundation models via reinforcement learning. Features M2Flow paradigm for embodied AI and agentic workflows with real-world robotics integrations. Apache 2.0 licensed.
a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more