User Guide

This section provides comprehensive guides for using TuFT in various scenarios.

Chat SFT

Supervised fine-tuning on chat-formatted data with assistant-only loss masking.

Chat Supervised Fine-Tuning (SFT)
Countdown RL

Reinforcement learning with GRPO-style training on verifiable tasks.

Countdown Reinforcement Learning (RL)
On-Policy Distillation

Distill a teacher into a student on the student’s own samples, via per-token reverse-KL.

On-Policy Distillation (OPD)
Custom Losses

Client-defined objectives (e.g. composite DPO + NLL) via forward_backward_custom, on HF and FSDP.

Custom Losses
LoRA Target Modules

How LoRA flags resolve to module names, and the rules enforced around them.

LoRA Target Modules
Persistence

Enable server state persistence with Redis for crash recovery.

Persistence
Observability

OpenTelemetry integration for tracing, metrics, and logs.

Observability (OpenTelemetry)
Console

Dashboard for monitoring training runs, checkpoints, and sampling playground.

User console