Skip to content
View fpolica91's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report fpolica91

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
fpolica91/README.md

Hi 👋 I'm Fabricio Policarpo

GPU Infrastructure Engineer — firmware-to-fabric

I commission, validate, and automate bare-metal GPU clusters — from PXE boot and BMC/Redfish, through GPU burn-in and NCCL/InfiniBand validation, up to inference serving. The work most teams never get to: catching lemon nodes before they kill a training run, and debugging the firmware / PCIe / fabric-level failures that don't show up in any dashboard.

A100 · H100 · B200 · RTX 5090


  • 🌍 South Florida · remote
  • 🔧 GPU bare-metal lifecycle — provisioning → validation → networking → serving
  • 📬 fabriciopolicarpo0@gmail.com

What I do

GPU cluster commissioning & validation — burn-in and acceptance suites (DCGM diag, gpu-burn, nccl-tests, sample training jobs), node health gating, lemon-node detection

Bare-metal automation — PXE / iPXE, Redfish / IPMI, Kea DHCP, NetBox, Ansible, cloud-init, Nomad — full lifecycle from rack to ready

Fabric & networking — InfiniBand / RoCE, ConnectX-7, rail-optimized topology, NCCL tuning, the kind of OEM misconfigs that get blamed on "the network"

Serving — vLLM, fp8/fp16 config, multi-tenant GPU provisioning on idle compute


Stack

Linux NVIDIA Ansible Terraform Nomad Python

Also: PXE/iPXE · Redfish/IPMI · Kea DHCP · NetBox · cloud-init · PostgreSQL · MongoDB · Node.js · NestJS · GraphQL · React · TypeScript


Featured

Chidori — GPU-over-IP platform. Remote GPU access over network fabric — low-level CUDA and systems work, TCP-tuned data path. A study in what it takes to move GPU workloads across the wire.

📝 Writing on GPU node burn-in, acceptance testing, and why LINPACK lies to you — in progress.


Find me

Pinned Loading

  1. chidori chidori Public

    GPU-over-IP: Transparent remote GPU access for CUDA compute workloads over TCP/IP

    C++ 1

  2. ALL_SCHOOL_42 ALL_SCHOOL_42 Public

    Forked from evgenkarlson/ALL_SCHOOL_42

    | SCHOOL_42_UPDATE 2020 | This repository contains ALL PROJECTS, TASKS AND SUBJECTS OF THE MAIN PROGRAM OF LEARNING AT SCHOOL 42 ( Program | Course | Programing | Coding | School 42 | Ecole 42 | Sc…

    C