Skip to content
Merged

Fixes #206

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
82 commits
Select commit Hold shift + click to select a range
f80b261
Register cellposev4 in benchmark run scripts
dariarom94 Jul 19, 2026
1a2fa09
fix anndata version mismatch with txsim
dariarom94 Jul 19, 2026
82add80
add segger to workflow (test)
dariarom94 Jul 19, 2026
53e1728
duplicates when FOV stiching cleaned up
dariarom94 Jul 19, 2026
1186b7a
chunks issue atera
dariarom94 Jul 20, 2026
18644d7
segger update image
dariarom94 Jul 20, 2026
ecb302d
claude fix for segger
dariarom94 Jul 20, 2026
7d66898
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
d400ebe
atera version fix
dariarom94 Jul 20, 2026
64d7b4e
wf for the custom rnaseq scripts
dariarom94 Jul 20, 2026
3edfbf1
adjust the loader image name
dariarom94 Jul 20, 2026
cbd2f12
adjust the memory
dariarom94 Jul 20, 2026
184260e
troubleshootig edges
dariarom94 Jul 20, 2026
9fa9a33
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
36631c4
segger update
dariarom94 Jul 21, 2026
0626127
cell type label correction
dariarom94 Jul 21, 2026
3186435
fix boundaries
dariarom94 Jul 21, 2026
d8a7d93
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
3505718
OOM fixes
dariarom94 Jul 21, 2026
d6e110a
fix code
dariarom94 Jul 21, 2026
4660f26
RCTD
dariarom94 Jul 21, 2026
5abd651
segger to RAPIDS
dariarom94 Jul 21, 2026
fe2e90a
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
0b23474
fix rctd
dariarom94 Jul 22, 2026
196ff1f
segger debug (torchvision)
dariarom94 Jul 22, 2026
4be7bd4
Merge branch 'main' into fixes
dariarom94 Jul 22, 2026
b8d3d7b
save the xenium version
dariarom94 Jul 22, 2026
202ac49
add atera to datasets
dariarom94 Jul 22, 2026
14be8d0
Add gene efficiency correction as a separate pipeline stage (#183)
dariarom94 Jul 22, 2026
0cf0243
moscot to pca and segger troubleshooting
dariarom94 Jul 22, 2026
d7afb84
added fastreseg
dariarom94 Jul 23, 2026
f87a1d9
segger bug new fix
dariarom94 Jul 23, 2026
123e112
fastreseg to workflow
dariarom94 Jul 23, 2026
7549589
add fastreseg test
dariarom94 Jul 23, 2026
9fa9604
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
ff04467
optimized fastreseg build
dariarom94 Jul 23, 2026
7ffc514
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
a7404d8
add s3 paths
dariarom94 Jul 23, 2026
19e5f83
troubleshoot comseg/segger
dariarom94 Jul 24, 2026
aaca151
segger update
dariarom94 Jul 25, 2026
c9bdb91
data loader bug
dariarom94 Jul 25, 2026
1df9834
Merge branch 'main' into fixes
dariarom94 Jul 25, 2026
4467d32
rctd adjustment (raw counts)
dariarom94 Jul 26, 2026
58912e4
fix segger and comseg
dariarom94 Jul 26, 2026
0a5aa99
optimize cosmx
dariarom94 Jul 26, 2026
298e666
Merge branch 'main' into fixes
dariarom94 Jul 26, 2026
c921937
parameter test for cellpose4
dariarom94 Jul 26, 2026
57c2d79
add atera
dariarom94 Jul 26, 2026
cf67e09
add a test in pciseq and dynamic memory for bruker
dariarom94 Jul 27, 2026
9f71692
add test to vizgen data
dariarom94 Jul 28, 2026
d795f33
Merge branch 'main' into fixes
dariarom94 Jul 28, 2026
fefaadc
param sweep
dariarom94 Jul 28, 2026
4dc08d3
add params to segmentation
dariarom94 Jul 28, 2026
e9505f1
adjust segger mem
dariarom94 Jul 29, 2026
3a155be
update fastreseg to tacco
dariarom94 Jul 29, 2026
acdd6c7
Add annotation + expression-correction parameter sweeps (rctd, ssam, …
dariarom94 Jul 30, 2026
fa462b8
Add moscot + split parameter sweeps (annotation, expression correction)
dariarom94 Jul 30, 2026
fd37aee
singler: read par['celltype_key'] instead of hardcoding "cell_type"
dariarom94 Jul 30, 2026
3449bd0
fastreseg
dariarom94 Jul 30, 2026
331d219
Merge branch 'main' into fixes
dariarom94 Jul 30, 2026
5961d0a
adjust labels
dariarom94 Jul 30, 2026
857e16a
bruker nsclc
dariarom94 Jul 31, 2026
54f023e
adjust bruker nsclc loader
dariarom94 Jul 31, 2026
6e8b9ea
setup
dariarom94 Jul 31, 2026
4eec493
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
44ad7af
method correction
dariarom94 Jul 31, 2026
b351148
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
2b8177c
pin anndata
dariarom94 Aug 1, 2026
6c16707
mirror nsclc
dariarom94 Aug 1, 2026
5618fb8
sync the vizgen files
dariarom94 Aug 2, 2026
b68c100
adjust mem for allen brain
dariarom94 Aug 2, 2026
98dbc69
claude notes
dariarom94 Aug 2, 2026
fa23294
claude notes
dariarom94 Aug 2, 2026
ac9492c
fix mirror script
dariarom94 Aug 2, 2026
26d292e
Merge branch 'main' into fixes
dariarom94 Aug 2, 2026
1ee371a
adapt fastreseg requirements
dariarom94 Aug 3, 2026
c53016a
merscope kuppe script update
dariarom94 Aug 3, 2026
d479db6
fix nsclc loader
dariarom94 Aug 3, 2026
6c54220
adjust mem for new test resources
dariarom94 Aug 4, 2026
9494fc5
adjust the nsclc loader for test resources
dariarom94 Aug 4, 2026
a459afc
change processor to avoid spatialdata 0.8.0 bug
dariarom94 Aug 4, 2026
8f7b0c1
pin spatialdata version
dariarom94 Aug 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 6 additions & 2 deletions scripts/create_resources/spatial/mirror_bruker_to_s3.sh
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,10 @@ SRC_BASE="https://smi-public.objects.liquidweb.services"
NSCLC_BASE="https://nanostring-public-share.s3.us-west-2.amazonaws.com/SMI-Compressed"
# Where the loaders read from (see the process_bruker_cosmx*.sh run scripts).
S3_DEST="${S3_DEST:-s3://openproblems-data/resources/raw_data/bruker_cosmx}"
# AWS CLI profile used for EVERY aws call. There is no default profile configured on the
# mirror host, so an aws call without this fails with NoCredentials -- which previously made
# the "already in S3?" size check silently return empty and re-download every file.
PROFILE="${PROFILE:-op}"
SCRATCH_DIR="${SCRATCH_DIR:-$PWD/bruker_mirror_scratch}"

# Files to mirror. Each entry is "SOURCE|LOCAL_NAME":
Expand Down Expand Up @@ -69,7 +73,7 @@ remote_size() {

s3_size() {
# Size of the object in S3, or empty if it doesn't exist
aws s3api head-object --bucket "$1" --key "$2" --query 'ContentLength' --output text 2>/dev/null || true
aws s3api head-object --profile "$PROFILE" --bucket "$1" --key "$2" --query 'ContentLength' --output text 2>/dev/null || true
}

# --- Main loop --------------------------------------------------------------
Expand Down Expand Up @@ -132,7 +136,7 @@ for entry in "${FILES[@]}"; do

# Upload to S3
echo " Uploading to s3://$bucket/$key ..."
aws s3 cp "$local_path" "s3://$bucket/$key" --profile op
aws s3 cp "$local_path" "s3://$bucket/$key" --profile "$PROFILE"

# Free scratch space before the next (much larger) file
echo " Removing local copy to free space"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ cd "$REPO_ROOT"

set -e

publish_dir="s3://openproblems-data/resources/datasets"
publish_dir="/scratch/task_ist_preprocessing/raw"

cat > /tmp/params.yaml << HERE
param_list:
Expand Down
12 changes: 12 additions & 0 deletions scripts/create_test_resources/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,15 @@ Download and process the spatial data:
`bash 2023_10x_mouse_brain_xenium_rep1.sh`

Combine the two datasets and run the ist preprocessing pipeline once with generic methods to create example outputs after each step: `test_pipeline.sh`

Two additional spatial test samples (cropped to a small region and published under `resources_test/common/`):
- 10x Atera WTA FFPE human breast cancer: `bash 2026_10x_human_breast_cancer_atera.sh`
- Bruker CosMx human lung cancer (Lung9 Rep1, smallest CosMx archive): `bash 2026_bruker_human_lung_cancer_cosmx.sh`

Each of the two also has a `*_nebius.sh` variant that runs on Nebius via `tw launch` and publishes to `/scratch/task_ist_preprocessing/resources_test/common/` instead of running locally (recommended for cosmx, which is too heavy for a laptop). After the run, sync the result from scratch to S3 (see the commented `aws s3 sync` at the bottom of each script).

Matching single-cell references for the two spatial samples (Nebius `tw launch` → scratch; each script runs the SC workflow, then chains the `subsample` module down to ~400 cells since these workflows don't subsample internally):
- Breast cancer reference for atera: `bash 2021_wu_human_breast_cancer_scrnaseq_nebius.sh`
- Lung cancer (NSCLC) reference for cosmx: `bash 2024_zuani_human_nsclc_scrnaseq_nebius.sh`

Pairings for a combined dataset (à la `mouse_brain_combined`): atera ↔ Wu breast, cosmx lung ↔ Zuani NSCLC.
5 changes: 4 additions & 1 deletion src/base/setup_spatialdata_partial.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
setup:
- type: python
pypi: ["spatialdata>=0.7.3", "anndata>=0.12.0,<0.13", "zarr>=3.0.0"]
# Pinned exactly: an unpinned >=0.7.3 silently pulled 0.8.0 on a rebuild, whose
# bounding_box/filter_table regressions transposed/emptied transcripts in the crop.
# process_dataset's transform-agnostic crop is validated against 0.8.0, so pin here.
pypi: ["spatialdata==0.8.0", "anndata>=0.12.0,<0.13", "zarr>=3.0.0"]
71 changes: 68 additions & 3 deletions src/data_processors/process_dataset/script.py
Original file line number Diff line number Diff line change
Expand Up @@ -75,9 +75,53 @@ def get_crop_coords(sdata, max_n_pixels=20000*20000): #50000*50000):
w_offset = (w - w_crop) // 2

crop = [[h_offset, h_offset + h_crop], [w_offset, w_offset + w_crop]]

return crop

def crop_points_by_global_xy(points, x0, x1, y0, y1):
"""Keep points whose GLOBAL (x, y) fall in [x0,x1] x [y0,y1], preserving the transform.

Global coords are computed from the element's affine matrix, so this is correct for
Scale, Affine (incl. reordered output axes like merscope's z,x,y) and 2D or 3D points
alike. It avoids spatialdata 0.8's bounding_box, which transposes Scale-transformed
transcripts and drops all points of a merscope Affine when the query omits z. Stays
lazy (dask) — no materialisation of large transcript tables.
"""
from spatialdata.models import PointsModel
from spatialdata.transformations import get_transformation, set_transformation

axes = [a for a in ("x", "y", "z") if a in points.columns]
trans = get_transformation(points, get_all=True)
# to_affine_matrix enforces "output axes superset of input axes" for Scale transforms,
# so request all axes as output and pick the x/y output rows.
M = trans["global"].to_affine_matrix(input_axes=tuple(axes), output_axes=tuple(axes))
ix, iy = axes.index("x"), axes.index("y")
gx = sum(M[ix, i] * points[a] for i, a in enumerate(axes)) + M[ix, len(axes)]
gy = sum(M[iy, i] * points[a] for i, a in enumerate(axes)) + M[iy, len(axes)]
keep = (gx >= x0) & (gx <= x1) & (gy >= y0) & (gy <= y1)
feat = points.attrs.get("spatialdata_attrs", {}).get("feature_key")
new = PointsModel.parse(points[keep], feature_key=feat) if feat else PointsModel.parse(points[keep])
set_transformation(new, trans, set_all=True)
return new

def crop_shapes_by_global_xy(shapes, x0, x1, y0, y1):
"""Keep shapes whose GLOBAL centroid falls in the box (mirrors the label pixel crop)."""
from shapely.affinity import affine_transform
from spatialdata.models import ShapesModel
from spatialdata.transformations import get_transformation, set_transformation

trans = get_transformation(shapes, get_all=True)
M = trans["global"].to_affine_matrix(input_axes=("x", "y"), output_axes=("x", "y"))
a, b, xoff, d, e, yoff = M[0, 0], M[0, 1], M[0, 2], M[1, 0], M[1, 1], M[1, 2]
c = shapes.geometry.apply(lambda g: affine_transform(g, [a, b, d, e, xoff, yoff]).centroid)
keep = c.map(lambda p: (x0 <= p.x <= x1) and (y0 <= p.y <= y1))
filtered = shapes[keep]
if len(filtered) == 0:
return None # nothing survives; spatialdata can't hold an empty shapes element -> drop it
new = ShapesModel.parse(filtered)
set_transformation(new, trans, set_all=True)
return new

def rechunk_sdata(sdata, CHUNK_SIZE=1024):
"""Rechunk the sdata to the given chunk size

Expand Down Expand Up @@ -252,16 +296,37 @@ def subsample_adata_group_balanced(adata, group_key, n_samples, seed=0):
sdata["image"] = sdata["morphology_mip"]
del sdata.images["morphology_mip"]

# spatialdata 0.8's filter_table join calls table.obs.reset_index(), which raises when a
# table's obs index name duplicates an obs column of the same name (merscope tables set
# EntityID as both the index and the instance_key column). Un-name the index to avoid the
# collision; the instance_key column is preserved, so the join still matches correctly.
for _tbl in sdata.tables.values():
if _tbl.obs.index.name is not None and _tbl.obs.index.name in _tbl.obs.columns:
_tbl.obs.index.name = None

# Crop datasets that are too large
crop_coords = get_crop_coords(sdata)
if crop_coords is not None:
(y0, y1), (x0, x1) = crop_coords # global box (raw image is at identity -> px == global)

# Raster elements (image, labels) + table crop correctly with bounding_box.
sdata_output = sdata.query.bounding_box(
axes=["y", "x"],
min_coordinate=[crop_coords[0][0], crop_coords[1][0]],
max_coordinate=[crop_coords[0][1], crop_coords[1][1]],
min_coordinate=[y0, x0],
max_coordinate=[y1, x1],
target_coordinate_system="global",
filter_table=True,
)
# spatialdata 0.8's bounding_box mis-crops TRANSFORMED vector elements: it transposes
# Scale-transformed transcripts (Xenium) and drops ALL points of a merscope Affine whose
# leading output axis is z when the query omits z (MERFISH). Re-crop points/shapes
# ourselves by their GLOBAL x/y — transform- and dimension-agnostic (see helpers above).
for _name in list(sdata.points):
sdata_output.points[_name] = crop_points_by_global_xy(sdata.points[_name], x0, x1, y0, y1)
for _name in list(sdata.shapes):
_cropped_shapes = crop_shapes_by_global_xy(sdata.shapes[_name], x0, x1, y0, y1)
if _cropped_shapes is not None: # None -> element lies entirely outside the crop
sdata_output.shapes[_name] = _cropped_shapes
# metadata is dataset-level, not spatial — re-add it if the bounding_box query dropped it
if "metadata" in sdata.tables and "metadata" not in sdata_output.tables:
sdata_output["metadata"] = sdata.tables["metadata"]
Expand Down
11 changes: 11 additions & 0 deletions src/datasets/loaders/bruker_cosmx/script.py
Original file line number Diff line number Diff line change
Expand Up @@ -477,6 +477,17 @@ def __getattr__(self, item):
def _fixed_global_cell_id(self, df):
return df["fov"] * (_RECON_MAX + 1) * (df["cell_ID"] > 0) + df["cell_ID"]
_cosmx_mod._CosMXReader._get_global_cell_id = _fixed_global_cell_id
else:
# For standard exports we have no precomputed max, but the same assertion trips whenever a
# segmented cell carries zero transcripts (tx-file max < exprMat/metadata max). read_tables
# runs before read_transcripts and tx cell_ID <= table cell_ID always, so reuse the running
# max across frames instead of asserting they agree.
def _running_max_global_cell_id(self, df):
max_cell_id = df["cell_ID"].max()
self.max_cell_id = max_cell_id if not hasattr(self, "max_cell_id") \
else max(self.max_cell_id, max_cell_id)
return df["fov"] * (self.max_cell_id + 1) * (df["cell_ID"] > 0) + df["cell_ID"]
_cosmx_mod._CosMXReader._get_global_cell_id = _running_max_global_cell_id

#########################################
# Convert raw files to spatialdata zarr #
Expand Down
20 changes: 20 additions & 0 deletions src/datasets/loaders/bruker_cosmx_nsclc/script.py
Original file line number Diff line number Diff line change
Expand Up @@ -245,6 +245,26 @@ def __getattr__(self, item):

_cosmx_mod.pd = _PandasReadCsvProxy(_real_pd)

############################################################
# Relax sopa's max-cell-id equality assertion (see below) #
############################################################
# sopa builds a collision-free global cell id as `fov * (max_cell_id + 1) + cell_ID`, where
# `max_cell_id` is the max *local* cell_ID across FOVs. It computes this multiplier the first
# time from the expression table (read_tables runs first) and then *asserts* the transcripts
# file has the exact same max. But a segmented cell can have zero detected transcripts, so the
# tx file's max local cell_ID can be smaller than the table's (Lung5_Rep2: table=5484, tx=4716)
# -> "Expected max cell ID to be 5484, but got 4716." The namespace only requires the SAME
# multiplier for both tables, not equal maxima; and a transcript's cell must exist in the
# segmentation, so tx cell_ID <= table cell_ID always. Reuse the running max instead of
# asserting so the table and transcripts share one multiplier.
def _patched_get_global_cell_id(self, df):
max_cell_id = df["cell_ID"].max()
self.max_cell_id = max_cell_id if not hasattr(self, "max_cell_id") \
else max(self.max_cell_id, max_cell_id)
return df["fov"] * (self.max_cell_id + 1) * (df["cell_ID"] > 0) + df["cell_ID"]

_cosmx_mod._CosMXReader._get_global_cell_id = _patched_get_global_cell_id

#########################################
# Convert raw files to spatialdata zarr #
#########################################
Expand Down
70 changes: 60 additions & 10 deletions src/datasets/loaders/vizgen_merscope/script.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# https://www.10xgenomics.com/datasets/fresh-frozen-mouse-brain-replicates-1-standard

import os
from copy import deepcopy
from datetime import datetime
from pathlib import Path

Expand Down Expand Up @@ -46,6 +47,22 @@
)


def _element_names(sd_obj):
"""List element names across spatialdata versions.

Newer spatialdata removed ``SpatialData.keys()``, so ``list(sdata.keys())`` raises
``AttributeError``. The per-type element dicts (``images``/``labels``/``points``/``shapes``/
``tables``) have been stable for far longer, so we collect names from those instead.
"""
names = []
for attr in ("images", "labels", "points", "shapes", "tables"):
try:
names.extend(getattr(sd_obj, attr).keys())
except Exception:
pass
return names


def _cell_polygon(cell_grp):
"""Boundary for one cell: prefer zIndex_3, else any available z-plane.

Expand Down Expand Up @@ -172,7 +189,7 @@ def read_boundary_hdf5(folder):
try:
print(
datetime.now() - t0,
"Loaded elements: " + ", ".join(list(sdata.keys())),
"Loaded elements: " + ", ".join(_element_names(sdata)),
flush=True,
)
except Exception as e:
Expand Down Expand Up @@ -202,15 +219,37 @@ def read_boundary_hdf5(folder):

print(datetime.now() - t0, "Renamed elements", flush=True)

# Rename transcript column
sdata["transcripts"] = sdata["transcripts"].rename(columns={"global_z": "z", "transcript_id": "ensembl_id"})#, "gene": "feature_name"})
if "gene" in sdata["transcripts"].columns:
# No idea why, but somehow dask dataframe renaming for the 'gene' column ends up in a key error when assigning it to sdata["transcripts"].
# update: see https://github.com/scverse/spatialdata/issues/996
sdata["transcripts"]["feature_name"] = sdata["transcripts"]["gene"]
del sdata["transcripts"]["gene"]
sdata['transcripts'].attrs["spatialdata_attrs"]["feature_key"] = "feature_name"
print(datetime.now() - t0, "Renamed transcripts column 'global_z' -> 'z' and 'gene' -> 'feature_name' and 'transcript_id' -> 'ensembl_id'", flush=True)
# Rename transcript columns.
# In newer pandas/dask, dask's DataFrame.rename() no longer propagates the frame's `.attrs`, so it
# returns a Points frame stripped of `transform` + `spatialdata_attrs`. Reassigning that bare frame
# to sdata["transcripts"] then fails PointsModel validation with
# "dask.dataframe.core.DataFrame.attrs does not contain `transform`" (scverse/spatialdata#996).
# We snapshot the valid `.attrs` before renaming and restore them afterwards, which reproduces the
# old (attrs-propagating) behaviour exactly and keeps the element model unchanged.
transcripts = sdata["transcripts"]
saved_attrs = deepcopy(dict(transcripts.attrs))

rename_map = {"global_z": "z", "transcript_id": "ensembl_id"}
renamed_gene = "gene" in transcripts.columns
if renamed_gene:
# The feature_key column ('gene') is renamed too; update the feature_key attr to match,
# otherwise validation KeyErrors on the now-missing 'gene' column.
rename_map["gene"] = "feature_name"
saved_attrs["spatialdata_attrs"]["feature_key"] = "feature_name"

renamed = transcripts.rename(columns=rename_map)
# Restore by mutating the frame's live `.attrs` dict in place (rather than reassigning `.attrs`),
# so this depends only on the getter returning the live dict -- the same assumption the rest of the
# loader already makes.
renamed.attrs.clear()
renamed.attrs.update(saved_attrs)
sdata["transcripts"] = renamed
print(
datetime.now() - t0,
"Renamed transcripts columns 'global_z' -> 'z', 'transcript_id' -> 'ensembl_id'"
+ (" and 'gene' -> 'feature_name'" if renamed_gene else ""),
flush=True,
)

print(datetime.now() - t0, "Columns in sdata['transcripts']:", sdata["transcripts"].columns, flush=True)

Expand Down Expand Up @@ -299,6 +338,17 @@ def read_boundary_hdf5(folder):

del sdata["cell_labels"].attrs["label_index_to_category"]

# Promote cell_labels from the single-scale DataArray that sd.rasterize() returns to a
# multiscale pyramid, matching the Xenium convention (vendor Xenium labels ship as a
# ["scale0"] DataTree). Without this the element is single-scale, and downstream
# components that index ["scale0"] on the segmentation copied from cell_labels
# (custom_segmentation) fail with KeyError: 'scale0'. scale_factors=[2,2,2,2] gives the
# scale0..scale4 pyramid; labels downsample with a label-preserving (nearest) method.
from spatialdata.models import Labels2DModel
# parse() preserves the element's existing transform; passing transformations= as well
# raises ("specified twice"), so don't.
sdata["cell_labels"] = Labels2DModel.parse(sdata["cell_labels"], scale_factors=[2, 2, 2, 2])


##############################
# Add info to metadata table #
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -79,4 +79,4 @@ runners:
- type: executable
- type: nextflow
directives:
label: [veryhighmem, midcpu, midtime]
label: [midmem, midcpu, midtime]
2 changes: 1 addition & 1 deletion src/datasets/loaders/zuani_human_nsclc_sc/config.vsh.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -75,4 +75,4 @@ runners:
- type: executable
- type: nextflow
directives:
label: [veryhighmem, midcpu, midtime]
label: [highmem, midcpu, midtime]
26 changes: 20 additions & 6 deletions src/datasets/loaders/zuani_human_nsclc_sc/script.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
## VIASH START

par = {
"input": "ftp://anonymous@ftp.ebi.ac.uk/biostudies/fire/E-MTAB-/526/E-MTAB-13526/Files/10X_Lung_Tumour_Annotated_v2.h5ad",
"input": "https://ftp.ebi.ac.uk/biostudies/fire/E-MTAB-/526/E-MTAB-13526/Files/10X_Lung_Tumour_Annotated_v2.h5ad",
"keep_files": True, # wether to delete the intermediate files
"output": "./temp/datasets/2024Zuani_human_nsclc_sc.h5ad",
"dataset_id": "2024Zuani_human_nsclc_sc",
Expand Down Expand Up @@ -36,11 +36,25 @@
# Resolving _viash_par (_viash_par)... failed: Name or service not known.
# wget: unable to resolve host address ‘_viash_par’)
FILE_PATH = TMP_DIR / "10X_Lung_Tumour_Annotated_v2.h5ad"
DOWNLOAD_URL = "ftp://anonymous@ftp.ebi.ac.uk/biostudies/fire/E-MTAB-/526/E-MTAB-13526/Files/10X_Lung_Tumour_Annotated_v2.h5ad"

# Download the data (55GB)
os.system(f'wget "{DOWNLOAD_URL}" -P "{TMP_DIR}/"')
# os.system(f'wget "{DOWNLOAD_URL}" -P "{TMP_DIR}/" --show-progress')
# EBI serves this file over HTTPS at the same host/path as the ftp:// URL. HTTPS avoids
# the FTP passive-mode data-connection timeouts ("PASV ... couldn't connect ... port
# NNNNN: Connection timed out") that make the ftp:// URL unreliable from cloud/cluster
# runners. The endpoint advertises Accept-Ranges: bytes, so `wget -c` resumes a partial
# download across the retries below.
DOWNLOAD_URL = "https://ftp.ebi.ac.uk/biostudies/fire/E-MTAB-/526/E-MTAB-13526/Files/10X_Lung_Tumour_Annotated_v2.h5ad"

# Download the data (~55GB). Resume-capable and retrying; check the exit code so a
# failed download raises a clear error here instead of a misleading FileNotFoundError
# when read_h5ad() is handed a path that was never written.
ret = os.system(
f'wget -c --tries=20 --timeout=120 --waitretry=60 --progress=dot:giga '
f'-O "{FILE_PATH}" "{DOWNLOAD_URL}"'
)
if ret != 0 or not FILE_PATH.exists():
raise RuntimeError(
f"Failed to download Zuani NSCLC h5ad from {DOWNLOAD_URL} "
f"(wget exit status {ret})."
)
adata = ad.read_h5ad(FILE_PATH)
# adata = adata[::100]

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ runners:
- type: executable
- type: nextflow
directives:
label: [ midtime, midcpu, midmem ]
label: [ veryhightime, midcpu, highmem ]


##macOS workaround: export DOCKER_DEFAULT_PLATFORM=linux/amd64
Loading