UHI-E3LLM is an end-to-end large language model (LLM)-based framework for forecasting monthly Urban Heat Island (UHI) intensity and generating factor-level explanations. The codebase provides the complete experimental workflows described in our paper, including data preprocessing, supervised fine-tuning (SFT), LoRA adapter fusion, PPO-based alignment, explainability enhancement, inference, baseline comparisons, leave-one-region-out evaluation, ablation studies, uncertainty quantification, and sensitivity analyses.
Important
For results that are consistent with the manuscript and for the best predictive performance, we strongly recommend DeepSeek-R1-Distill-Llama-8B. The smaller DeepSeek-R1-Distill-Qwen-1.5B may be used on resource-limited machines to check installation and the inference workflow, but its outputs are not expected to reproduce the manuscript results or match the performance of the 8B model.
UHI_E3LLM/
|-- baseline/ # Four tabular and three time-series baselines
|-- data/ # Versioned city manifests; downloaded datasets are ignored
|-- docs/ # Inputs, splits, and manuscript-to-code map
|-- environment/ # Unified locked environment for all workflows
|-- scripts/ # Executable manuscript workflows
|-- src/ # UHI-E3LLM training, inference, reward, and parsing
|-- LICENSE
|-- README.md
`-- requirements.txt
Detailed documentation:
The experiments were conducted using the environment listed below.
| Component | Experimental environment |
|---|---|
| Operating system | Ubuntu 20.04.6 LTS, Linux kernel 5.15.0 |
| CPU | Intel Xeon Platinum 8368 @ 2.40 GHz |
| System RAM | 1.0 TiB |
| GPU | NVIDIA A100-SXM4-80GB |
| NVIDIA driver | 575.57.08 |
| CUDA version | CUDA 11.8 |
| Python | 3.10 |
conda env create -f environment/environment.yml
conda activate uhi-e3llmA representative normal desktop configuration for the resource-limited demo is:
- 64-bit Linux, or Windows 11 with WSL2;
- an 8-core or better desktop CPU;
- 32 GB system RAM;
- an NVIDIA GeForce RTX 5060 Ti with 16 GB VRAM;
- at least 50 GB of free SSD space.
This configuration is intended for running the DeepSeek-R1-Distill-Qwen-1.5B workflow. It can be used to verify installation, model loading, inference, and output generation. For results consistent with the manuscript, we strongly recommend using DeepSeek-R1-Distill-Llama-8B.
The complete processed dataset is published under CC BY 4.0:
- DOI: https://doi.org/10.5281/zenodo.22161925
- File:
Dataset.zip
Download Dataset.zip from Zenodo and extract the archive. Place the extracted
train/, valid/, city_metrics/, regional_folds/, and
explanation_samples/ directories directly under this repository's data/
directory. The required layout is documented in data/README.md.
huggingface-cli download deepseek-ai/DeepSeek-R1-Distill-Llama-8B \
--local-dir model/deepseek-8b \
--local-dir-use-symlinks FalseThe trained LoRA adapters are available from Kurofu/UHI_E3LLM.
Download the adapter collection into the code-repository root:
huggingface-cli download Kurofu/UHI_E3LLM \
--local-dir Deepseek8b_Adapters \
--local-dir-use-symlinks FalseThe released artifacts are organized as follows:
Deepseek8b_Adapters/
|-- sft/checkpoint-epoch0/ ... checkpoint-epoch4/
|-- SFT_fusion/
|-- ppo/epoch0/
|-- ppo/epoch1/
|-- distillation/epoch0/
|-- distillation/epoch1/
`-- run_all_stage_inference.sh
After downloading the processed dataset and base model, run the evaluation script with:
conda activate uhi-e3llm
GPU_ID=0 bash Deepseek8b_Adapters/run_all_stage_inference.sh- environment creation and Python-package installation: approximately 10–30 minutes on a typical broadband-connected desktop;
- repository download: usually less than 5 minutes;
- model download: approximately 30–120 minutes, depending on the selected model, network bandwidth, connection stability, and Hugging Face server availability. The actual download time may vary substantially.
conda activate uhi-e3llm
bash scripts/run_sft.sh
bash scripts/run_lora_fusion.sh
bash scripts/run_ppo.sh
bash scripts/run_explanation_distillation.sh
bash scripts/run_main_evaluation.sh
bash scripts/run_efficiency_measurement.shWith the default paths, the complete pipeline produces the following main artifacts. Individual output roots can be changed through the environment variables accepted by the corresponding workflow scripts.
data/explanation_samples/
|-- sample_1.txt
|-- ...
`-- sample_900.txt
model_adapters/
|-- SFT/deepseek-8b/
| |-- checkpoint-epoch0/
| |-- ...
| `-- checkpoint-epoch4/
|-- merged/deepseek-8b/
|-- PPO/deepseek-8b_merged__Global500/
| |-- epoch0/
| `-- epoch1/
`-- explanation_distilled/deepseek-8b_merged__Global500/
|-- epoch0/
`-- epoch1/
model/
|-- deepseek-8b_merged__Global500/
|-- deepseek-8b_merged__Global500_PPO_epoch0_merged/
`-- deepseek-8b_merged__Global500_PPO_epoch1_merged/
outputs/
|-- lora_fusion/
|-- ppo_train_baseline/
|-- main_ensemble/
| |-- epoch0/
| |-- epoch1/
| `-- ensemble/
|-- baselines/
|-- efficiency/
| |-- raw/
| `-- efficiency_measurements.csv
`-- tables/
|-- table1_main_performance.csv
|-- table2_leave_one_region_out.csv
`-- table3_ablation.csv
logs/
`-- <workflow-name>.log
bash scripts/run_leave_one_region_out.shbash scripts/run_module_ablations.sh
bash scripts/run_component_ablations.shbash scripts/run_city_bootstrap.sh
bash scripts/run_independent_calibration.shbash scripts/run_ppo_penalty_sensitivity.sh
bash scripts/run_lora_similarity_threshold.sh# Table 1: main test-set and city-size comparison
bash scripts/run_tables.sh table1
# Table 2: leave-one-region-out comparison
bash scripts/run_tables.sh table2
# Table 3: module- and component-level ablations
bash scripts/run_tables.sh table3
# Rebuild all three after all experiments are complete
bash scripts/run_tables.sh allTable 1 requires run_main_evaluation.sh; Table 2 requires
run_leave_one_region_out.sh; Table 3 requires the main evaluation and both
ablation workflows.
The ablation results are reported in Table 3 and are not rendered as a
separate bar-chart panel by scripts/run_figures.sh.
The plotting workflow covers the manuscript panels generated directly from experimental outputs. Fig. 3(c--e), Fig. 5(b,d), and Fig. 6(b--d) were manually composed using HTML, while Fig. 2(d--f) and Fig. 4(a) were manually prepared in QGIS.
GPU_ID=0 bash scripts/run_case_study_outputs.sh
GPU_ID=0 bash scripts/run_main_evaluation.sh
GPU_ID=0 bash scripts/run_leave_one_region_out.sh
GPU_ID=0 bash scripts/run_module_ablations.sh
GPU_ID=0 bash scripts/run_component_ablations.sh
bash scripts/run_city_bootstrap.sh
bash scripts/run_independent_calibration.sh
bash scripts/run_ppo_penalty_sensitivity.sh
EFFICIENCY_RUN_MAIN_EVALUATION=0 bash scripts/run_efficiency_measurement.sh
bash scripts/run_figures.shThe released data/explanation_samples/sample_*.txt files are consumed
directly by the default distillation workflow and do not require API access.
Regenerating those samples, or constructing fold-specific explanation data,
uses the DeepSeek API. Never store an API key in source code:
export DEEPSEEK_API_KEY="your-key"API access is not required for numerical inference when the model and adapter checkpoints are available locally.
The following estimates assume the resource-limited demo configuration: an NVIDIA GeForce RTX 5060 Ti with 16 GB VRAM, 32 GB system RAM, an 8-core or better CPU, and SSD storage. The estimates apply to the 1.5B workflow and are provided for planning purposes rather than as measured benchmarks.
| Stage | Estimated runtime |
|---|---|
| SFT | 8–14 h |
| LoRA adapter merging | 1–5 min |
| PPO training | 28–48 h |
| Explanation-text distillation | 10–30 min |
| Inference and explanation generation | 1–3 h |
UHI-E3LLM is released under the MIT License.