Fine-tuning method that merges the SFT and preference-optimisation stages, further simplifying the DPO pipeline.