Skip to content
View kennethdsheridan's full-sized avatar

Block or report kennethdsheridan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
kennethdsheridan/README.md

Kenny Sheridan

I build software and design infrastructure systems that help founders, ML researchers, and customer engineering teams run GPU training and inference workloads on production infrastructure.

After eight years as a Meteorologist in the U.S. Marine Corps, I moved into systems engineering across bare metal, GPU platforms, cloud infrastructure, and research environments. For the last several years, I’ve worked in fast-moving startups turning ambiguous infrastructure and product requirements into shipped, reliable systems.

My work spans customer workload onboarding, distributed training, batch inference, ML pipeline observability, autoscaling infrastructure, utilization-based billing, and workload resiliency. I've written and deployed custom Kubernetes controllers, host agents, and control-plane integrations, and help adapt manifests, containers, serving stacks, model caches, telemetry, and deployment workflows so customer workloads run reliably across virtual machines, Kubernetes, and bare-metal GPU fleets.

I have hands-on production experience with NVIDIA Hopper/Blackwell and AMD Instinct environments, vLLM, SGLang, TensorRT-LLM, Triton, PyTorch, Candle, tch-rs, NCCL debugging, kernel and BIOS tuning, high-bandwidth data movement, and large-scale GPU performance engineering.

My infrastructure work spans Kubernetes, Slurm on Kubernetes (SUNK), Software Defined Networking, AWS, GCP, DigitalOcean, co-located bare metal, Terraform/OpenTofu, Terragrunt, Argo CD, Cluster API, Kamaji, QEMU/KVM, Prometheus/Grafana, Weka, VAST, Ceph, Postgres, Google Cloud SQL, Rust, Go, and Nix/NixOS.

Recent work includes multi-cluster observability for 12,000+ GPUs across 10+ providers and regions, customer workload onboarding for training and inference, hosted metrics, automated Slack alerting, Stripe-backed utilization billing, Cluster API-compatible infrastructure APIs, custom Kubernetes controllers, host-side agents, control-plane automation, and reproducible test/dev environments using Nix, Tilt, Kubernetes, and QEMU/KVM.

I’m also an advocate for local inference, with hands-on experience with tools for M1/M2 Apple Silicon using Asahi Linux diagnostics, and exploring Apple Container orchestration with MLX and JACCL.

Links:

Pinned Loading

  1. sfcompute/hardware_report sfcompute/hardware_report Public

    Produces a serialized hardware report of the physical infrastructure for automation

    Rust 29 2

  2. sysperf-svr sysperf-svr Public

    System benchmarking, profiling and stressing

    Rust 1

  3. nvim nvim Public

    Lua