Skip to content

Repository files navigation

hprofiler — Heterogeneous Profiler

Multi-device CPU/GPU profiler for Linux. Traces programs across CUDA, ROCm, OpenCL, OpenMP, NCCL, and MPI simultaneously — with a terminal UI (and an optional native Qt GUI, see GUI Viewer below), a live Flame Graph tab and a native TUI roofline viewer, and cross-layer causal attribution: one dependency graph over every backend active in a run, with a confidence-graded, formally-computed critical path instead of a single-runtime or heuristic one (see Cross-Layer Causal Attribution below). CPU sampling is provided via Linux perf.

Requirements

  • Python 3.10+, CMake 3.16+, GCC/Clang
  • pip install click textual rich capstone
  • TUI roofline viewer: pip install plotly "kaleido==0.2.1" (0.2.1 specifically — later versions require Chrome and break on clusters); the Flame Graph tab needs neither
  • Backend-specific: CUDA toolkit, a libamdhip64 (ROCm/HIP) runtime, LLVM libomp, an MPI implementation (mpicc, or a Cray Programming Environment cc wrapper), or perf — see DOCUMENTATION.md for exact search paths

Build

pip install click textual rich capstone
python3 hprofiler build

Produces build/lib/libhprofiler_{cuda,opencl,ompt,rocm,nccl,mpi}.so.

Quick Start

# Profile with auto-detected backends
python3 hprofiler run -- ./my_program

# Specific backends
python3 hprofiler run --backend cuda,cpu      -- ./cuda_app
python3 hprofiler run --backend openmp        -- ./omp_app
python3 hprofiler run --backend rocm          -- ./hip_app
python3 hprofiler run --backend cuda,nccl     -- ./multi_gpu_app
python3 hprofiler run --backend mpi           -- mpirun -np 4 ./mpi_app

# Call tree (adds Call Tree tab; compile with -fno-omit-frame-pointer -rdynamic)
python3 hprofiler run --call-tree --backend cuda -- ./app

# Per-kernel disassembly (adds Disasm tab)
python3 hprofiler run --backend cuda --disasm -- ./app

# Instruction-level GPU heat map + stall annotation (CUDA, libcupti.so loaded at runtime)
python3 hprofiler run --backend cuda --disasm --gpu-pc-sampling -- ./app

# Instruction-level CPU heat (OpenCL CPU runtime via ACPP)
ACPP_VISIBILITY_MASK=ocl python3 hprofiler run --backend opencl,cpu --disasm -- ./app

# Save trace, skip TUI
python3 hprofiler run --no-ui -o trace.json -- ./app

# Open a saved trace
python3 hprofiler view trace.json

# Text summary only
python3 hprofiler summary trace.json

# Native Qt GUI instead of the TUI (falls back to the TUI automatically
# if PySide6/X11 aren't available — see GUI Viewer below)
python3 hprofiler run --gui --backend cuda -- ./cuda_app
python3 hprofiler gui trace.json

# Flame graph — populates the Flame Graph tab in both the TUI and the GUI
python3 hprofiler run --perf-callgraph dwarf -- ./my_program
python3 hprofiler run --backend cuda --perf-callgraph dwarf --call-tree -- ./cuda_app  # + GPU API overhead

# Roofline chart — opens TUI viewer by default (requires plotly + kaleido)
python3 hprofiler roofline --backend cuda    -- ./cuda_app
python3 hprofiler roofline --backend cpu     -- ./cpu_app
python3 hprofiler roofline --backend rocm    -- ./hip_app
python3 hprofiler roofline --html --backend cuda -- ./cuda_app  # write HTML + open browser

# Hardware PMU counters via LIKWID
HPROFILER_LIKWID_GROUP=MEM python3 hprofiler run --backend likwid -- ./app

# List available backends on this machine
python3 hprofiler backends

Always separate hprofiler options from the target program with --.

Backends

Name Alias Injection What is traced
cpu perf perf record subprocess CPU samples, optional DWARF/fp/lbr call-graph
cuda — LD_PRELOAD Kernel launches, memcpy, syncs, NVTX ranges, memory counters
opencl cl LD_PRELOAD Kernel enqueues, buffer transfers, JIT compile time
openmp omp OMP_TOOL_LIBRARIES (OMPT) Parallel regions, tasks, loops, barriers (requires LLVM libomp, not GCC libgomp)
rocm hip LD_PRELOAD HIP kernel launches, memcpy, memory counters
nccl — LD_PRELOAD Collectives (AllReduce, Broadcast, …), point-to-point — GPU-accurate timing
mpi — PMPI / LD_PRELOAD Send/Recv, collectives, one-sided ops — wall-clock timing
likwid hwc likwid-perfctr wrapper Hardware PMU counters: FLOPS, DRAM bandwidth, cache rates, CPI

TUI Viewer

Opens automatically after hprofiler run. Tabs:

Tab When shown Contents
System Always Device specs, FP16/32/64/Tensor TFLOP/s, bandwidth, IPC, LLC/branch miss rates, RSS
Profile Always GPU activity%, time breakdown by category, top hotspots, bottleneck advisor
Timeline Always Gantt view with per-stream CUDA/ROCm lanes
Hotspots Always Filterable/sortable function table
Call Tree Only with --call-tree and/or --perf-callgraph fp|dwarf|lbr Stack-frame tree from captured call graphs
Flame Graph Same condition as Call Tree Proportional icicle chart of the same call-stack data — see Flame Graph Tab Controls
Disasm Only with --disasm Per-kernel assembly with instruction-type color coding, runtime heat % and stall columns (CPU via perf, CUDA via --gpu-pc-sampling), and static optimization hints

GUI Viewer

An optional native Qt/QML desktop GUI (pip install "hprofiler[gui]") covering the same tabs as the TUI (including Flame Graph), plus a few GUI-specific additions: smooth wheel-zoom/drag-pan on the Timeline, a scrollable node-and-edge call-graph panel showing which functions call which for whatever's currently visible, an idle-time overlay, so a span blocked at a nested barrier/sync call visibly shows that within its own bar instead of looking continuously busy, and a 3-panel Source tab with instruction-mix/static-advisor analysis alongside the assembly.

hprofiler run --gui --backend cuda -- ./cuda_app
hprofiler gui trace.hprofiler.json
hprofiler run --gui --perf-callgraph dwarf -- ./app   # + populate the Flame Graph tab

Falls back to the TUI automatically — no error shown — if PySide6 isn't installed, X11 isn't reachable, or GPU-rendered Qt Quick fails over indirect/forwarded X11 (retried once with software rendering first). See DOCUMENTATION.md §21 for the full tab reference and §2 for install/troubleshooting (including the libxcb-cursor0 system-library requirement and a VNC fallback for machines where installing it isn't an option).

Flame Graph Tab Controls

Works in any terminal — plain character-cell rendering, no inline-image protocol required (unlike the roofline viewer below). Same controls in the GUI's Flame Graph tab, mouse-driven there too.

Action TUI GUI
Zoom into a frame Click Click
Zoom out one level Right-click or Backspace Right-click
Reset to full view Escape Escape / Reset button
Search — regex-highlight matching frames Type in the search box Type in the search box

Roofline TUI Controls

Key Action
n / p Cycle through kernels (shows crosshairs with headroom annotation)
Esc Deselect kernel / hide crosshairs
+ / = Zoom in
- Zoom out
← → ↑ ↓ Pan
r Reset zoom
w Open HTML version in browser
q Quit

Output Files

File Viewer
<prog>.hprofiler.json Perfetto or chrome://tracing
<prog>.roofline.html Any browser (self-contained)

OpenTelemetry Export

Export spans and metrics to any OTLP-compatible collector. No extra Python dependencies — uses stdlib only.

# Send live to a local collector (Grafana Alloy, otelcol, Jaeger ≥ 1.35, Tempo, …)
python3 hprofiler run --backend cuda --otlp-endpoint http://localhost:4318 -- ./app

# Write OTLP JSON to file (replay later with curl)
python3 hprofiler run --backend cuda --otlp-file trace.otlp.json -- ./app

# Export from a saved trace
python3 hprofiler view --otlp-endpoint http://localhost:4318 app.hprofiler.json
python3 hprofiler view --otlp-file trace.otlp.json app.hprofiler.json

# Replay a saved OTLP file to a collector
curl -X POST http://localhost:4318/v1/traces \
     -H 'Content-Type: application/json' -d @trace.otlp.json

OTLP mapping: each SpanEvent becomes an OTLP span (all root-level, no parent inference); CounterEvent values (IPC, bandwidth, memory usage) become OTLP gauge metrics sent to /v1/metrics; hprofiler category, tags, PID, and TID become span attributes.

Disassembly

Pass --disasm to collect post-run per-kernel disassembly (runs in background, TUI opens immediately):

Backend Tool needed
CUDA AoT cuobjdump (CUDA toolkit)
CUDA JIT (ACPP) Built-in PTX parser
ROCm llvm-objdump (apt install llvm)
CPU / OpenMP capstone (pip install capstone) or objdump
OpenCL JIT (ACPP SSCP generic) objdump on the .jit.so emitted by ACPP SSCP
OpenCL CPU (Intel CPU OCL) objdump on x86-64 ELF extracted via clGetProgramInfo

Cross-Layer Causal Attribution

hprofiler's core contribution: for programs combining several backends at once (e.g. MPI+OpenMP+CUDA), it builds one dependency graph over every captured span — CUDA, ROCm, OpenCL, OpenMP, MPI, NCCL together, not a separate per-runtime trace to merge — using each programming model's real synchronization semantics (resolved MPI wildcard matching, real communicator identity, stream/device-sync ordering, OpenMP barriers), and finds the true critical path via a formal DAG longest-path computation, not a heuristic walk. Every edge in that graph is tagged with how directly the underlying data proves it (certain/high/medium — see DOCUMENTATION.md §18), so the result never presents a call-order guess with the same confidence as a hardware-enforced ordering.

# POP-style parallel efficiency breakdown (Load Balance, Communication
# Efficiency, GPU/NCCL efficiency, ...) computed from a single trace
python3 hprofiler efficiency trace.json

# N-way cross-runtime critical path + blame attribution across every
# backend active in the trace (generalizes CASITA/HPCToolkit-style
# critical-path analysis beyond MPI+CUDA-only or CPU+GPU-only), with a
# per-hop confidence breakdown
python3 hprofiler critical-path trace.json

# Multi-node: profile each node separately, then merge onto one timeline
# before running critical-path/efficiency across node boundaries
python3 hprofiler merge-nodes node0.json node1.json node2.json -o merged.json
python3 hprofiler critical-path merged.json

See DOCUMENTATION.md §17–20 for the exact formulas, edge-confidence model, formal critical-path algorithm, multi-node clock synchronization, and what's approximate vs. exact vs. still hardware-unverified on this development machine (documented honestly, not glossed over — see §13's Known Limitations table).

See DOCUMENTATION.md for the full CLI reference, backend details, wire protocol, and how to extend the profiler.

AI Performance Analysis

hprofiler can use an LLM to analyse a profile and produce a written report of bottlenecks, root causes, and prioritized optimization recommendations.

Quick start

# Analyse an existing trace (auto-detects LLM from env vars)
python3 hprofiler analyze trace.hprofiler.json

# Profile and analyse in one step
python3 hprofiler analyze --backend cuda -- ./app

# Add AI analysis to the normal run workflow
python3 hprofiler run --analyze --backend cuda -- ./app

# Compare two runs and report what changed
python3 hprofiler analyze --compare before.json after.json trace.hprofiler.json

# Save report to a Markdown file
python3 hprofiler analyze --output-report report.md trace.hprofiler.json

LLM provider setup

The provider is auto-detected from environment variables. Set one of the following before running:

# Anthropic Claude (recommended)
export ANTHROPIC_API_KEY=sk-ant-...
python3 hprofiler analyze trace.hprofiler.json
# Defaults to claude-sonnet-4-6; override with --llm-model claude-opus-4-8

# OpenAI GPT
export OPENAI_API_KEY=sk-...
python3 hprofiler analyze trace.hprofiler.json
# Defaults to gpt-4o

# Ollama (local, no API key needed)
ollama serve                        # start the Ollama daemon
ollama pull llama3.1:8b             # pull a model
python3 hprofiler analyze trace.hprofiler.json
# Defaults to llama3.1:8b; any pulled model works

# Any OpenAI-compatible endpoint (vLLM, LM Studio, Groq, Together.ai, …)
python3 hprofiler analyze \
  --llm openai-compat \
  --llm-endpoint http://localhost:8080 \
  --llm-model Qwen2.5-72B-Instruct \
  trace.hprofiler.json

Persistent configuration via environment variables

export HPROFILER_LLM_PROVIDER=anthropic   # anthropic | openai | ollama | openai-compat
export HPROFILER_LLM_MODEL=claude-opus-4-8
export HPROFILER_LLM_API_KEY=sk-ant-...   # if not using ANTHROPIC_API_KEY / OPENAI_API_KEY
export HPROFILER_LLM_ENDPOINT=http://...  # for openai-compat / custom Ollama host

How it works

The agent calls a suite of tools to drill into the profile — hotspots, kernel details, memory patterns, timeline phases, synchronisation overhead, MPI communication — before writing its final report. For models without tool-use support it falls back to a single comprehensive prompt. No new Python packages are required; all HTTP calls use urllib.request.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages