Fine-tuning / sft
Supervised fine-tuning
The plain approach: continue training on curated demonstrations. It remains the first thing to try, the thing preference methods are initialised from, and the setting in which most claims about parameter-efficient methods are actually evaluated.
1 entry.