Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Preference Tuning Under Domain Shift

Official code for An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

Constantinos Karouzos · Xingwei Tan · Nikolaos Aletras

Accepted at EMNLP 2026 Main Conference arXiv 2601.05882 Project website Python 3.10 or newer MIT License

Paper · Project website · Citation

Study design for preference tuning under domain shift: adaptation strategies feed into five alignment objectives and are evaluated for generalization and diversity.

Navigate

Quick start · Repository map · Data · Training · Pseudo-labeling · Generation and evaluation · Citation

Repository map

Path Contents
configs/templates/ Runnable paper-level configurations for SFT, DPO, KTO, ORPO, RM, PPO, and GRPO
prefadap/training/ Training argument schemas, pipelines, and shared training utilities
prefadap/pseudo_label/ Local-teacher pseudo-label generation and DPO/KTO output formats
prefadap/runtime/ Generation helpers and portable vLLM policies
prefadap/evaluation/ Diversity and safety evaluation implementations
prefadap/cli/ Training, generation, pseudo-labeling, judging, and moderation entry points
scripts/ Dataset preparation utilities

Quick start

Python 3.10 or newer is recommended.

git clone https://github.com/ckarouzos/prefadap.git
cd prefadap
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Pseudo-labeling and high-throughput generation require a compatible vLLM installation. Install vLLM separately in the CUDA or ROCm environment where it will run.

Check the installation and available objectives with:

python -m prefadap.cli.main list pipelines

Data

The public dataset identifiers are:

Testbed Source Target
Summarization tldr cnndm
Helpfulness shp_askengineers shp_askculinary
Safety pku_saferlhf_cybercrime pku_saferlhf_violence_physicalharm

TL;DR and CNN/DailyMail are loaded through Hugging Face Datasets. SHP domain files are expected at data/prefadapt_data/SHP/<domain>/<split>.jsonl; records must retain history, human_ref_A, human_ref_B, and labels.

PKU-SafeRLHF can be streamed directly from Hugging Face and filtered without a separate download or a full local copy:

python scripts/filter_pku_domains.py \
  --domain cybercrime \
  --split train \
  --output data/pku_saferlhf/cybercrime/train.jsonl

python scripts/filter_pku_domains.py \
  --domain violence_physicalharm \
  --split train \
  --output data/pku_saferlhf/violence_physicalharm/train.jsonl

Repeat with --split test and a test.jsonl output path to prepare evaluation data. For offline use, pass --input path/to/local.jsonl; the same domain filter is then applied to the local file. The loaders preserve existing chosen/rejected orientation and otherwise place the safer upstream response in chosen.

Training

Paper-level training defaults are provided in configs/templates/.

The same entry point supports sft, dpo, kto, orpo, rm, ppo, and grpo. The templates use one epoch, a maximum source length of 1024, LoRA rank 16 with alpha 256 and dropout 0.05, cosine scheduling with 2% warmup, and bfloat16. Offline methods use a maximum target length of 512; PPO and GRPO use 128-token responses. Safety experiments use four epochs. The Llama-3.1-8B method defaults are:

Method Micro-batch / accumulation Learning rate Objective settings
SFT 8 / 16 1e-5 —
DPO 4 / 32 1e-5 beta: 0.1
KTO 8 / 16 8e-7 beta: 0.1, equal desirable/undesirable weights
ORPO 4 / 32 1e-5 beta: 0.1
RM 8 / 16 1e-5 —
PPO 64 / 1 1e-6 KL 0.1, clip 0.1, value coefficient 0.05
GRPO 64 / 1 1e-6 group size 4, beta: 0.1, epsilon: 0.2

For OLMo-3-7B SFT, the paper uses a micro-batch of 4, accumulation of 8, and learning rate 8e-5.

PPO and GRPO templates require reward_model_name_or_path. Multi-GPU runs should be launched with Accelerate and a site-appropriate DeepSpeed ZeRO-3 configuration.

Pseudo-labeling

The teacher pipeline uses meta-llama/Llama-3.3-70B-Instruct, three samples per input, temperature 0.7, top-p 0.9, and a maximum of 512 new tokens. Its input is JSONL with prompt and chosen fields, where chosen is the available target-domain reference. The prompt family is selected explicitly.

python -m prefadap.cli.pseudo_label_pipeline \
  --dataset-path data/cnndm_prompts.jsonl \
  --dataset-family summarisation \
  --dataset-name cnndm \
  --output-format dpo \
  --run-id cnndm_teacher \
  --output data/pseudo_cnndm.jsonl \
  --tensor-parallel-size 4

--dataset-family selects the task prompt, while --dataset-name selects dataset-specific context limits. If --dataset-name is omitted, the target dataset for that family is used. Use --output-format kto for unary desirable/undesirable examples. For training, set dataset_name to the corresponding pseudo_* identifier and pseudo_data_path to the generated JSONL.

Generation and evaluation

Generate held-out outputs with the paper decoding configuration:

RUN_DIR="$PWD" python -m prefadap.cli.generate \
  --model outputs/dpo_cnndm/final_model \
  --output-dir outputs/generations/dpo_cnndm \
  --decoding-config configs/decoding.yaml \
  --datasets cnndm \
  --backend hf

For vLLM, use a merged Hugging Face checkpoint and --backend vllm.

Pairwise judging

The judge utility creates OpenAI Batch API request files without making network calls. It uses gpt-5-nano-2025-08-07, alternates response positions, and writes a manifest used to undo position swaps during scoring.

python -m prefadap.cli.nano_batch_judge prepare \
  --candidate-a outputs/generations/base/cnndm_test.jsonl \
  --candidate-b outputs/generations/dpo/cnndm_test.jsonl \
  --label-a base --label-b dpo --dataset cnndm \
  --output outputs/judge/requests.jsonl

python -m prefadap.cli.nano_batch_judge score \
  --results outputs/judge/results.jsonl \
  --manifest outputs/judge/requests.manifest.jsonl \
  --output outputs/judge/summary.json

Add --domain-aware for the AskCulinary domain-aware rubric reported in the paper. Submit the prepared JSONL and download the result JSONL with the OpenAI Batch API separately.

Safety

Safety scoring follows the same offline pattern with omni-moderation-latest:

python -m prefadap.cli.evaluate_moderation prepare \
  --generations outputs/generations/safety.jsonl \
  --output outputs/moderation/requests.jsonl

python -m prefadap.cli.evaluate_moderation score \
  --results outputs/moderation/results.jsonl \
  --manifest outputs/moderation/requests.manifest.jsonl \
  --output outputs/moderation/summary.json

Diversity

The paper protocol samples 16 generations for each of 500 prompts at temperature 1.0 and reports EAD, Sentence-BERT, and NLI diversity.

python -m prefadap.cli.evaluate_diversity outputs/generations/diversity \
  --max-inputs 500 \
  --generations-per-input 16 \
  --temperature 1.0 \
  --output outputs/diversity.json

The generation settings are recorded in configs/diversity/diversity_protocol.yaml.

Citation

@misc{karouzos2026preference,
  title         = {An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift},
  author        = {Constantinos Karouzos and Xingwei Tan and Nikolaos Aletras},
  year          = {2026},
  eprint        = {2601.05882},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}

Acknowledgments

We thank Jasivan Sivakumar, Vynska Amalia Permadi, Sam Lewis-Lim, Yanwen Peng, and Atsuki Yamaguchi for their help and feedback. CK is supported by the UKRI Centre for Doctoral Training in Speech and Language Technologies and their Applications under grant EP/S023062/1. XT and NA are supported by the EPSRC under grant EP/Y009800/1 through Responsible AI UK Keystone Project KP0016.

We acknowledge the University of Sheffield IT Services for high-performance computing services and the University of Oxford Advanced Research Computing facility. We also acknowledge the EuroHPC Joint Undertaking for access to the LEONARDO supercomputer, hosted by CINECA and the LEONARDO consortium through a EuroHPC Development Access call. This work used resources from the Isambard-AI National AI Research Resource, operated by the University of Bristol and funded by the UK Government's Department for Science, Innovation and Technology through UK Research and Innovation and the Science and Technology Facilities Council under grant ST/AIRR/I-A-I/1023.

License

The code is released under the MIT License. Upstream datasets, models, and hosted services retain their own terms.

About

Official code for An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Contributors

Languages