Multi-device CPU/GPU profiler for Linux. Traces programs across CUDA, ROCm, OpenCL, OpenMP, NCCL, and MPI simultaneously — with a terminal UI (and an optional native Qt GUI, see GUI Viewer below), a live Flame Graph tab and a native TUI roofline viewer, and cross-layer causal attribution: one dependency graph over every backend active in a run, with a confidence-graded, formally-computed critical path instead of a single-runtime or heuristic one (see Cross-Layer Causal Attribution below). CPU sampling is provided via Linux perf.
- Python 3.10+, CMake 3.16+, GCC/Clang
pip install click textual rich capstone- TUI roofline viewer:
pip install plotly "kaleido==0.2.1"(0.2.1 specifically — later versions require Chrome and break on clusters); the Flame Graph tab needs neither - Backend-specific: CUDA toolkit, a
libamdhip64(ROCm/HIP) runtime, LLVMlibomp, an MPI implementation (mpicc, or a Cray Programming Environmentccwrapper), orperf— see DOCUMENTATION.md for exact search paths
pip install click textual rich capstone
python3 hprofiler buildProduces build/lib/libhprofiler_{cuda,opencl,ompt,rocm,nccl,mpi}.so.
# Profile with auto-detected backends
python3 hprofiler run -- ./my_program
# Specific backends
python3 hprofiler run --backend cuda,cpu -- ./cuda_app
python3 hprofiler run --backend openmp -- ./omp_app
python3 hprofiler run --backend rocm -- ./hip_app
python3 hprofiler run --backend cuda,nccl -- ./multi_gpu_app
python3 hprofiler run --backend mpi -- mpirun -np 4 ./mpi_app
# Call tree (adds Call Tree tab; compile with -fno-omit-frame-pointer -rdynamic)
python3 hprofiler run --call-tree --backend cuda -- ./app
# Per-kernel disassembly (adds Disasm tab)
python3 hprofiler run --backend cuda --disasm -- ./app
# Instruction-level GPU heat map + stall annotation (CUDA, libcupti.so loaded at runtime)
python3 hprofiler run --backend cuda --disasm --gpu-pc-sampling -- ./app
# Instruction-level CPU heat (OpenCL CPU runtime via ACPP)
ACPP_VISIBILITY_MASK=ocl python3 hprofiler run --backend opencl,cpu --disasm -- ./app
# Save trace, skip TUI
python3 hprofiler run --no-ui -o trace.json -- ./app
# Open a saved trace
python3 hprofiler view trace.json
# Text summary only
python3 hprofiler summary trace.json
# Native Qt GUI instead of the TUI (falls back to the TUI automatically
# if PySide6/X11 aren't available — see GUI Viewer below)
python3 hprofiler run --gui --backend cuda -- ./cuda_app
python3 hprofiler gui trace.json
# Flame graph — populates the Flame Graph tab in both the TUI and the GUI
python3 hprofiler run --perf-callgraph dwarf -- ./my_program
python3 hprofiler run --backend cuda --perf-callgraph dwarf --call-tree -- ./cuda_app # + GPU API overhead
# Roofline chart — opens TUI viewer by default (requires plotly + kaleido)
python3 hprofiler roofline --backend cuda -- ./cuda_app
python3 hprofiler roofline --backend cpu -- ./cpu_app
python3 hprofiler roofline --backend rocm -- ./hip_app
python3 hprofiler roofline --html --backend cuda -- ./cuda_app # write HTML + open browser
# Hardware PMU counters via LIKWID
HPROFILER_LIKWID_GROUP=MEM python3 hprofiler run --backend likwid -- ./app
# List available backends on this machine
python3 hprofiler backendsAlways separate hprofiler options from the target program with --.
| Name | Alias | Injection | What is traced |
|---|---|---|---|
cpu |
perf |
perf record subprocess |
CPU samples, optional DWARF/fp/lbr call-graph |
cuda |
— | LD_PRELOAD | Kernel launches, memcpy, syncs, NVTX ranges, memory counters |
opencl |
cl |
LD_PRELOAD | Kernel enqueues, buffer transfers, JIT compile time |
openmp |
omp |
OMP_TOOL_LIBRARIES (OMPT) |
Parallel regions, tasks, loops, barriers (requires LLVM libomp, not GCC libgomp) |
rocm |
hip |
LD_PRELOAD | HIP kernel launches, memcpy, memory counters |
nccl |
— | LD_PRELOAD | Collectives (AllReduce, Broadcast, …), point-to-point — GPU-accurate timing |
mpi |
— | PMPI / LD_PRELOAD | Send/Recv, collectives, one-sided ops — wall-clock timing |
likwid |
hwc |
likwid-perfctr wrapper |
Hardware PMU counters: FLOPS, DRAM bandwidth, cache rates, CPI |
Opens automatically after hprofiler run. Tabs:
| Tab | When shown | Contents |
|---|---|---|
| System | Always | Device specs, FP16/32/64/Tensor TFLOP/s, bandwidth, IPC, LLC/branch miss rates, RSS |
| Profile | Always | GPU activity%, time breakdown by category, top hotspots, bottleneck advisor |
| Timeline | Always | Gantt view with per-stream CUDA/ROCm lanes |
| Hotspots | Always | Filterable/sortable function table |
| Call Tree | Only with --call-tree and/or --perf-callgraph fp|dwarf|lbr |
Stack-frame tree from captured call graphs |
| Flame Graph | Same condition as Call Tree | Proportional icicle chart of the same call-stack data — see Flame Graph Tab Controls |
| Disasm | Only with --disasm |
Per-kernel assembly with instruction-type color coding, runtime heat % and stall columns (CPU via perf, CUDA via --gpu-pc-sampling), and static optimization hints |
An optional native Qt/QML desktop GUI (pip install "hprofiler[gui]") covering the same tabs as the TUI (including Flame Graph), plus a few GUI-specific additions: smooth wheel-zoom/drag-pan on the Timeline, a scrollable node-and-edge call-graph panel showing which functions call which for whatever's currently visible, an idle-time overlay, so a span blocked at a nested barrier/sync call visibly shows that within its own bar instead of looking continuously busy, and a 3-panel Source tab with instruction-mix/static-advisor analysis alongside the assembly.
hprofiler run --gui --backend cuda -- ./cuda_app
hprofiler gui trace.hprofiler.json
hprofiler run --gui --perf-callgraph dwarf -- ./app # + populate the Flame Graph tabFalls back to the TUI automatically — no error shown — if PySide6 isn't installed, X11 isn't reachable, or GPU-rendered Qt Quick fails over indirect/forwarded X11 (retried once with software rendering first). See DOCUMENTATION.md §21 for the full tab reference and §2 for install/troubleshooting (including the libxcb-cursor0 system-library requirement and a VNC fallback for machines where installing it isn't an option).
Works in any terminal — plain character-cell rendering, no inline-image protocol required (unlike the roofline viewer below). Same controls in the GUI's Flame Graph tab, mouse-driven there too.
| Action | TUI | GUI |
|---|---|---|
| Zoom into a frame | Click | Click |
| Zoom out one level | Right-click or Backspace | Right-click |
| Reset to full view | Escape | Escape / Reset button |
| Search — regex-highlight matching frames | Type in the search box | Type in the search box |
| Key | Action |
|---|---|
n / p |
Cycle through kernels (shows crosshairs with headroom annotation) |
| Esc | Deselect kernel / hide crosshairs |
+ / = |
Zoom in |
- |
Zoom out |
← → ↑ ↓ |
Pan |
r |
Reset zoom |
w |
Open HTML version in browser |
q |
Quit |
| File | Viewer |
|---|---|
<prog>.hprofiler.json |
Perfetto or chrome://tracing |
<prog>.roofline.html |
Any browser (self-contained) |
Export spans and metrics to any OTLP-compatible collector. No extra Python dependencies — uses stdlib only.
# Send live to a local collector (Grafana Alloy, otelcol, Jaeger ≥ 1.35, Tempo, …)
python3 hprofiler run --backend cuda --otlp-endpoint http://localhost:4318 -- ./app
# Write OTLP JSON to file (replay later with curl)
python3 hprofiler run --backend cuda --otlp-file trace.otlp.json -- ./app
# Export from a saved trace
python3 hprofiler view --otlp-endpoint http://localhost:4318 app.hprofiler.json
python3 hprofiler view --otlp-file trace.otlp.json app.hprofiler.json
# Replay a saved OTLP file to a collector
curl -X POST http://localhost:4318/v1/traces \
-H 'Content-Type: application/json' -d @trace.otlp.jsonOTLP mapping: each SpanEvent becomes an OTLP span (all root-level, no parent inference); CounterEvent values (IPC, bandwidth, memory usage) become OTLP gauge metrics sent to /v1/metrics; hprofiler category, tags, PID, and TID become span attributes.
Pass --disasm to collect post-run per-kernel disassembly (runs in background, TUI opens immediately):
| Backend | Tool needed |
|---|---|
| CUDA AoT | cuobjdump (CUDA toolkit) |
| CUDA JIT (ACPP) | Built-in PTX parser |
| ROCm | llvm-objdump (apt install llvm) |
| CPU / OpenMP | capstone (pip install capstone) or objdump |
| OpenCL JIT (ACPP SSCP generic) | objdump on the .jit.so emitted by ACPP SSCP |
| OpenCL CPU (Intel CPU OCL) | objdump on x86-64 ELF extracted via clGetProgramInfo |
hprofiler's core contribution: for programs combining several backends at
once (e.g. MPI+OpenMP+CUDA), it builds one dependency graph over every
captured span — CUDA, ROCm, OpenCL, OpenMP, MPI, NCCL together, not a
separate per-runtime trace to merge — using each programming model's real
synchronization semantics (resolved MPI wildcard matching, real
communicator identity, stream/device-sync ordering, OpenMP barriers), and
finds the true critical path via a formal DAG longest-path computation,
not a heuristic walk. Every edge in that graph is tagged with how directly
the underlying data proves it (certain/high/medium — see
DOCUMENTATION.md §18), so the result never presents a
call-order guess with the same confidence as a hardware-enforced ordering.
# POP-style parallel efficiency breakdown (Load Balance, Communication
# Efficiency, GPU/NCCL efficiency, ...) computed from a single trace
python3 hprofiler efficiency trace.json
# N-way cross-runtime critical path + blame attribution across every
# backend active in the trace (generalizes CASITA/HPCToolkit-style
# critical-path analysis beyond MPI+CUDA-only or CPU+GPU-only), with a
# per-hop confidence breakdown
python3 hprofiler critical-path trace.json
# Multi-node: profile each node separately, then merge onto one timeline
# before running critical-path/efficiency across node boundaries
python3 hprofiler merge-nodes node0.json node1.json node2.json -o merged.json
python3 hprofiler critical-path merged.jsonSee DOCUMENTATION.md §17–20 for the exact formulas, edge-confidence model, formal critical-path algorithm, multi-node clock synchronization, and what's approximate vs. exact vs. still hardware-unverified on this development machine (documented honestly, not glossed over — see §13's Known Limitations table).
See DOCUMENTATION.md for the full CLI reference, backend details, wire protocol, and how to extend the profiler.
hprofiler can use an LLM to analyse a profile and produce a written report of bottlenecks, root causes, and prioritized optimization recommendations.
# Analyse an existing trace (auto-detects LLM from env vars)
python3 hprofiler analyze trace.hprofiler.json
# Profile and analyse in one step
python3 hprofiler analyze --backend cuda -- ./app
# Add AI analysis to the normal run workflow
python3 hprofiler run --analyze --backend cuda -- ./app
# Compare two runs and report what changed
python3 hprofiler analyze --compare before.json after.json trace.hprofiler.json
# Save report to a Markdown file
python3 hprofiler analyze --output-report report.md trace.hprofiler.jsonThe provider is auto-detected from environment variables. Set one of the following before running:
# Anthropic Claude (recommended)
export ANTHROPIC_API_KEY=sk-ant-...
python3 hprofiler analyze trace.hprofiler.json
# Defaults to claude-sonnet-4-6; override with --llm-model claude-opus-4-8
# OpenAI GPT
export OPENAI_API_KEY=sk-...
python3 hprofiler analyze trace.hprofiler.json
# Defaults to gpt-4o
# Ollama (local, no API key needed)
ollama serve # start the Ollama daemon
ollama pull llama3.1:8b # pull a model
python3 hprofiler analyze trace.hprofiler.json
# Defaults to llama3.1:8b; any pulled model works
# Any OpenAI-compatible endpoint (vLLM, LM Studio, Groq, Together.ai, …)
python3 hprofiler analyze \
--llm openai-compat \
--llm-endpoint http://localhost:8080 \
--llm-model Qwen2.5-72B-Instruct \
trace.hprofiler.jsonexport HPROFILER_LLM_PROVIDER=anthropic # anthropic | openai | ollama | openai-compat
export HPROFILER_LLM_MODEL=claude-opus-4-8
export HPROFILER_LLM_API_KEY=sk-ant-... # if not using ANTHROPIC_API_KEY / OPENAI_API_KEY
export HPROFILER_LLM_ENDPOINT=http://... # for openai-compat / custom Ollama hostThe agent calls a suite of tools to drill into the profile — hotspots, kernel details, memory patterns, timeline phases, synchronisation overhead, MPI communication — before writing its final report. For models without tool-use support it falls back to a single comprehensive prompt. No new Python packages are required; all HTTP calls use urllib.request.