A collection of small, runnable implementations of LLM post-training and alignment methods, from RL basics to DPO, RLHF, and RLAIF.
reinforcement-learning gymnasium featured ppo reinforcement-learning-agent policy-based-method value-based-methods grpo status-learning
-
Updated
Sep 20, 2026 - Python