Skip to content
WB-PIDA-Data-Science-ShopPublic

About

A package designed for replicating the ETL (Extract, Transform, Load) of Country-Level Institutional Assessment and Review (CLIAR) data

Resources

Stars

1 star

Watchers

1 watching

Forks

Repository files navigation

cliaretl

R-CMD-check

The goal of cliaretl is to be the central data pipeline responsible for sourcing, processing and delivering the data to the CLIAR dashboard. Its primary function is to execute a robust Extract, Transform and Load (ETL) process, ensuring the dashboard always has access to accurate and harmonized data. It provides the following features:

  • Automated Data Extraction: Programmatically pulls raw data from various external services through efficient API calls.

  • Data Harmonization: Merges and standardizes data originating from both API endpoints and manual input sources. This ensures data consistency and reliability, regardless of its origin.

  • Data Preparation: Transforms raw and harmonized data into a clean, structured, and optimized format, making it available to use in the CLIAR dashboard. This process also facilitates version control for the data, ensuring traceability and reproducibility.

  • Data Quality Testing: Implements a modular testing approach to verify data quality at various stages of the ETL pipeline, ensuring data integrity and reliability before it reaches the dashboard.

Project Structure

The Package is structured to facilitate the ETL process, with the following key components:

cliaretl/
├── .github/             # GitHub configuration files (e.g., actions, workflows).
├── analysis/            # Data Transformation and quality control scripts.
├── data/                # Clean indicators datasets ready to be Loaded.
├── data-raw/            # Data extraction scripts.
│   ├── input/           # Original raw data (from external sources).
│   ├── source/          # Scripts to transform raw data.
│   └── output/          # Intermediate datasets for checks or ad-hoc use.
├── inst/                # Package resources (e.g., extdata, templates).
├── man/                 # Auto-generated R documentation (via roxygen2).
├── R/                   # Core R functions of the package.
├── renv/                # R environment and dependency management (via renv).
├── spielplatz/          # Sandbox/experimentation area.
├── tests/               # Unit tests and testthat framework scripts.
├── .gitignore           # Files and folders ignored by Git.
├── .Rbuildignore        # Files excluded from package build.
├── .Renviron            # Environment variables for the project.
├── .Rprofile            # Project-specific R startup settings.
├── cliaretl.Rproj       # RStudio project file.
├── DESCRIPTION          # R package metadata and dependencies.
├── LICENSE              # License (standard CRAN format).
├── LICENSE.md           # License (Markdown format).
├── NAMESPACE            # Package namespace (functions to export/import).
├── README.md            # Project documentation (rendered version).
├── README.Rmd           # Source file for README.md.
├── renv.lock            # Lockfile for reproducible package versions.

Package Installation

Install the cliaretl package with:

# Step 1. Install the packages 'pak' or 'remotes' if you don't have them:

# install.packages("pak")
# install.packages("remotes")


# Step 2. Install cliaretl from GitHub:
remotes::install_github("WB-PIDA-Data-Science-Shop/cliaretl")

Reproducibility

The cliaretl package uses renv to lock package versions, ensuring consistent results across environments and dependencies management. To set up the project environment for package development, run renv::restore().

About

A package designed for replicating the ETL (Extract, Transform, Load) of Country-Level Institutional Assessment and Review (CLIAR) data

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages