User Guide¶
This section provides comprehensive guides for using TuFT in various scenarios.
Chat SFT
Supervised fine-tuning on chat-formatted data with assistant-only loss masking.
Countdown RL
Reinforcement learning with GRPO-style training on verifiable tasks.
On-Policy Distillation
Distill a teacher into a student on the student’s own samples, via per-token reverse-KL.
Custom Losses
Client-defined objectives (e.g. composite DPO + NLL) via forward_backward_custom, on HF and FSDP.
LoRA Target Modules
How LoRA flags resolve to module names, and the rules enforced around them.
Persistence
Enable server state persistence with Redis for crash recovery.
Observability
OpenTelemetry integration for tracing, metrics, and logs.
Console
Dashboard for monitoring training runs, checkpoints, and sampling playground.