Official code for An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
Quick start · Repository map · Data · Training · Pseudo-labeling · Generation and evaluation · Citation
| Path | Contents |
|---|---|
configs/templates/ |
Runnable paper-level configurations for SFT, DPO, KTO, ORPO, RM, PPO, and GRPO |
prefadap/training/ |
Training argument schemas, pipelines, and shared training utilities |
prefadap/pseudo_label/ |
Local-teacher pseudo-label generation and DPO/KTO output formats |
prefadap/runtime/ |
Generation helpers and portable vLLM policies |
prefadap/evaluation/ |
Diversity and safety evaluation implementations |
prefadap/cli/ |
Training, generation, pseudo-labeling, judging, and moderation entry points |
scripts/ |
Dataset preparation utilities |
Python 3.10 or newer is recommended.
git clone https://github.com/ckarouzos/prefadap.git
cd prefadap
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtPseudo-labeling and high-throughput generation require a compatible vLLM installation. Install vLLM separately in the CUDA or ROCm environment where it will run.
Check the installation and available objectives with:
python -m prefadap.cli.main list pipelinesThe public dataset identifiers are:
| Testbed | Source | Target |
|---|---|---|
| Summarization | tldr |
cnndm |
| Helpfulness | shp_askengineers |
shp_askculinary |
| Safety | pku_saferlhf_cybercrime |
pku_saferlhf_violence_physicalharm |
TL;DR and CNN/DailyMail are loaded through Hugging Face Datasets. SHP domain files are expected at data/prefadapt_data/SHP/<domain>/<split>.jsonl; records must retain history, human_ref_A, human_ref_B, and labels.
PKU-SafeRLHF can be streamed directly from Hugging Face and filtered without a separate download or a full local copy:
python scripts/filter_pku_domains.py \
--domain cybercrime \
--split train \
--output data/pku_saferlhf/cybercrime/train.jsonl
python scripts/filter_pku_domains.py \
--domain violence_physicalharm \
--split train \
--output data/pku_saferlhf/violence_physicalharm/train.jsonlRepeat with --split test and a test.jsonl output path to prepare evaluation data. For offline use, pass --input path/to/local.jsonl; the same domain filter is then applied to the local file. The loaders preserve existing chosen/rejected orientation and otherwise place the safer upstream response in chosen.
Paper-level training defaults are provided in configs/templates/.
The same entry point supports sft, dpo, kto, orpo, rm, ppo, and grpo. The templates use one epoch, a maximum source length of 1024, LoRA rank 16 with alpha 256 and dropout 0.05, cosine scheduling with 2% warmup, and bfloat16. Offline methods use a maximum target length of 512; PPO and GRPO use 128-token responses. Safety experiments use four epochs. The Llama-3.1-8B method defaults are:
| Method | Micro-batch / accumulation | Learning rate | Objective settings |
|---|---|---|---|
| SFT | 8 / 16 | 1e-5 |
— |
| DPO | 4 / 32 | 1e-5 |
beta: 0.1 |
| KTO | 8 / 16 | 8e-7 |
beta: 0.1, equal desirable/undesirable weights |
| ORPO | 4 / 32 | 1e-5 |
beta: 0.1 |
| RM | 8 / 16 | 1e-5 |
— |
| PPO | 64 / 1 | 1e-6 |
KL 0.1, clip 0.1, value coefficient 0.05 |
| GRPO | 64 / 1 | 1e-6 |
group size 4, beta: 0.1, epsilon: 0.2 |
For OLMo-3-7B SFT, the paper uses a micro-batch of 4, accumulation of 8, and learning rate 8e-5.
PPO and GRPO templates require reward_model_name_or_path. Multi-GPU runs should be launched with Accelerate and a site-appropriate DeepSpeed ZeRO-3 configuration.
The teacher pipeline uses meta-llama/Llama-3.3-70B-Instruct, three samples per input, temperature 0.7, top-p 0.9, and a maximum of 512 new tokens. Its input is JSONL with prompt and chosen fields, where chosen is the available target-domain reference. The prompt family is selected explicitly.
python -m prefadap.cli.pseudo_label_pipeline \
--dataset-path data/cnndm_prompts.jsonl \
--dataset-family summarisation \
--dataset-name cnndm \
--output-format dpo \
--run-id cnndm_teacher \
--output data/pseudo_cnndm.jsonl \
--tensor-parallel-size 4--dataset-family selects the task prompt, while --dataset-name selects dataset-specific context limits. If --dataset-name is omitted, the target dataset for that family is used. Use --output-format kto for unary desirable/undesirable examples. For training, set dataset_name to the corresponding pseudo_* identifier and pseudo_data_path to the generated JSONL.
Generate held-out outputs with the paper decoding configuration:
RUN_DIR="$PWD" python -m prefadap.cli.generate \
--model outputs/dpo_cnndm/final_model \
--output-dir outputs/generations/dpo_cnndm \
--decoding-config configs/decoding.yaml \
--datasets cnndm \
--backend hfFor vLLM, use a merged Hugging Face checkpoint and --backend vllm.
The judge utility creates OpenAI Batch API request files without making network calls. It uses gpt-5-nano-2025-08-07, alternates response positions, and writes a manifest used to undo position swaps during scoring.
python -m prefadap.cli.nano_batch_judge prepare \
--candidate-a outputs/generations/base/cnndm_test.jsonl \
--candidate-b outputs/generations/dpo/cnndm_test.jsonl \
--label-a base --label-b dpo --dataset cnndm \
--output outputs/judge/requests.jsonl
python -m prefadap.cli.nano_batch_judge score \
--results outputs/judge/results.jsonl \
--manifest outputs/judge/requests.manifest.jsonl \
--output outputs/judge/summary.jsonAdd --domain-aware for the AskCulinary domain-aware rubric reported in the paper. Submit the prepared JSONL and download the result JSONL with the OpenAI Batch API separately.
Safety scoring follows the same offline pattern with omni-moderation-latest:
python -m prefadap.cli.evaluate_moderation prepare \
--generations outputs/generations/safety.jsonl \
--output outputs/moderation/requests.jsonl
python -m prefadap.cli.evaluate_moderation score \
--results outputs/moderation/results.jsonl \
--manifest outputs/moderation/requests.manifest.jsonl \
--output outputs/moderation/summary.jsonThe paper protocol samples 16 generations for each of 500 prompts at temperature 1.0 and reports EAD, Sentence-BERT, and NLI diversity.
python -m prefadap.cli.evaluate_diversity outputs/generations/diversity \
--max-inputs 500 \
--generations-per-input 16 \
--temperature 1.0 \
--output outputs/diversity.jsonThe generation settings are recorded in configs/diversity/diversity_protocol.yaml.
@misc{karouzos2026preference,
title = {An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift},
author = {Constantinos Karouzos and Xingwei Tan and Nikolaos Aletras},
year = {2026},
eprint = {2601.05882},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}We thank Jasivan Sivakumar, Vynska Amalia Permadi, Sam Lewis-Lim, Yanwen Peng, and Atsuki Yamaguchi for their help and feedback. CK is supported by the UKRI Centre for Doctoral Training in Speech and Language Technologies and their Applications under grant EP/S023062/1. XT and NA are supported by the EPSRC under grant EP/Y009800/1 through Responsible AI UK Keystone Project KP0016.
We acknowledge the University of Sheffield IT Services for high-performance computing services and the University of Oxford Advanced Research Computing facility. We also acknowledge the EuroHPC Joint Undertaking for access to the LEONARDO supercomputer, hosted by CINECA and the LEONARDO consortium through a EuroHPC Development Access call. This work used resources from the Isambard-AI National AI Research Resource, operated by the University of Bristol and funded by the UK Government's Department for Science, Innovation and Technology through UK Research and Innovation and the Science and Technology Facilities Council under grant ST/AIRR/I-A-I/1023.
The code is released under the MIT License. Upstream datasets, models, and hosted services retain their own terms.