WHY THIS
EXISTS.
Fine-tuning a local model still feels like cluster work, while many experiments get promoted because one demo looked good. The missing piece is a repeatable path from curated examples to a measured, reversible adapter.
// Make the score earn trust
A local training and evaluation workbench for small language models, LoRA adapters, human feedback and measurable promotion gates.
Fine-tuning a local model still feels like cluster work, while many experiments get promoted because one demo looked good. The missing piece is a repeatable path from curated examples to a measured, reversible adapter.
LLM Gym keeps training pools, LoRA definitions, hardware-aware backends, verification and promotion in one local dashboard. Adapters move forward only after explicit checks and human gates.
Define one narrow skill, curate its JSONL pool and train a small swappable LoRA instead of another monolith.
MLX, PEFT and simulate backends share a single queue that avoids overlapping jobs and memory blowups.
Acceptance prompts, local judges, pool forecasts and DPO feedback turn promotion into a measurable decision.