I run MarMar Labs, an independent software and research lab in Minneapolis, Minnesota. I build AI products, contribute fixes to the open-source tools I use, and do security research through HackerOne.
My work sits where AI agents, developer tools, and real-world reliability meet: making useful software, tracing what breaks, and backing up a fix with evidence.
Website · HackerOne · X · LinkedIn · Email
| Project | What I'm building |
|---|---|
| SignAI | An iPhone app for scanning, understanding, signing, and tracking agreements. Optional AI review, an agreement vault, and deadline reminders. On the App Store. |
| stui | Streamlit-inspired interfaces that run in the terminal. Python widgets, session state, reruns, tables, and a full-screen data explorer, built on Textual. Install from PyPI. |
| RetryProof | Deterministic retry-fault testing for n8n workflows. Human-approved invariants, AI-proposed repairs, and before/after replay evidence in a controlled lab. Try the lab. |
| NeverGuess | Preflight for AI-assisted code changes: architecture, risks, missing tests, rollout notes, and a better prompt before the agent starts editing. |
| AgentBridge | An experimental action layer for agent-ready apps: typed manifests, permissioned MCP tools, confirmations, and audit logs. TypeScript SDK. |
| MarSWE-Bench | An agentic coding benchmark with original multi-file debugging tasks across five languages, hidden behavioral tests, and a public score/cost leaderboard. |
155 merged pull requests across five upstream projects, as of October 9, 2026. These are contributions to other maintainers' repositories, with a few examples below.
Browse my public merged pull requests →
I research trust boundaries in developer tools, AI agents, APIs, CLIs, and CI/CD systems. My approach combines source review, controlled reproductions, responsible disclosure, and retesting fixes.
My published security record documents $7,000 in rewards, 13 distinct rewarded reports, and 7 retest awards. That is the September 11, 2026 public snapshot; the reward total includes retest awards.
Find me on HackerOne as realmarmarlabs. I share public outcomes and methodology while keeping private reports and program details confidential.
Frontier AI stress-test audit — a reproducibility-oriented collection of prompts, transcripts, artifacts, and verification notes on reasoning and code generation. An artifact-based audit, with its limitations documented. Read the preprint.
I work across Python, TypeScript/JavaScript, Go, React Native, and developer infrastructure. I use AI coding tools heavily, then check the result against the actual application: reproduce the failure, make the smallest useful fix, and verify the behavior.
If you're building agent tooling, practical AI products, or open-source infrastructure, reach me at founder@marmarlabs.com.



