I went to school for graphic design because I loved art. Drawing, digital work, making things by hand and on-screen. That was always where I felt most like myself. I graduated from FAMU and almost immediately realized that making art for other people was not the path for me. Handing creative work over to a client and watching it become something safer, flatter, and more generic was demoralizing in a way I was not prepared for. So I pivoted. I took a tech support role at Dropbox, and while I did not love the calls, I loved the learning. I loved becoming a subject matter expert, understanding new systems quickly, and knowing what was coming before most people did. That feeling has never really left me.
The pattern across every role since has been fairly consistent. I come in early, find the gaps, and build. At Facebook, I started as a QA contractor on a team of 12. I built the onboarding process, wrote documentation that outlasted my time there, and was functioning as an unofficial team lead by the time the team had grown to 50. That work led to a full-time project manager role on a different team and, as it turned out, my entry into ML operations. From there, I was recruited to Pinterest as the first hire on their data labeling and relevance program manager team. At The Browser Company, I built their LLM evaluation framework from scratch and stood up a data pipeline to analyze user feedback at scale. The throughline is clear to me now: I do my best work when things are new, messy, and undefined. I like building the system, setting the standard, and leaving things in better shape than I found them.
Also, the graphic design background never went away. It shows up in every presentation I build, every tool I design, every document I put my name on. I used to treat it as a side skill, something adjacent to the real work. Over time I've realized I'm exercising the same muscle. I want everything I build to work well and look like it was made with intention.
Choose an area to explore
Human Data Operations is the infrastructure layer behind model quality, safety, and product reliability, even when it gets treated like back-office work. I've spent eight years building programs that hold up under pressure: annotation systems from scratch at Meta and Pinterest, eight-figure vendor portfolios, and 20 million-plus annotations a year across search, ads, and trust and safety.
Vendor portfolio and operations. I've managed vendor portfolios north of $16M, run full RFP processes, and coordinated workforces of 300-plus across dozens of locales. I know how to build accountability into vendor relationships early, so quality, cost, and delivery issues are visible before they become emergencies.
Annotation infrastructure and tooling. I've owned annotation platform selection, implementation, and technical coordination with engineering, including full RFPs, vendor migrations, workflow design, and integration requirements. At Pinterest, I completed a dual-platform transition in six months against a typical one-year timeline.
QA and quality infrastructure. I build annotation QA systems from the ground up, including IRR monitoring, calibration workflows, sampling protocols, and escalation paths. At Pinterest, the QA infrastructure I designed improved annotation quality 20 percent across all programs.
Safety and responsible AI programs. I've supported ML safety classifier programs at Meta and built trust and safety annotation infrastructure at Pinterest, including EAP support for raters reviewing sensitive content. When Pinterest needed to mobilize quickly on a content issue, the operational structure was already in place.
ML lifecycle. I treat annotation as part of the model performance feedback loop, not a standalone production function. That means surfacing re-training data needs, identifying annotation gaps tied to model regressions, and connecting what models struggle with back to the labeling pipelines that can improve them.
Cross-functional influence. I translate data quality into business risk so engineering, product, and leadership teams understand why the infrastructure matters. At Pinterest, the QA program I built started as self-funded informal audits and became a formal investment with dedicated headcount.
Direct management and mentorship. I've directly managed and trained QA specialists, contractors, and program managers across Meta and Pinterest. As the founding PM on Pinterest's data labeling and relevance team, I trained every program manager who joined after me and built the documentation, onboarding, and operating practices that kept the program from depending on one person.
Tools: Appen / Figure Eight, Labelbox, Google Sheets + Apps Script, BigQuery, SQL, Tableau, Retool
LLM evaluation only works when a team can define quality before trying to measure it. At The Browser Company, I built the eval program from zero as the sole person responsible for both the framework design and the measurement infrastructure, turning subjective model behavior into signal that product and engineering teams could actually use.
Eval framework design. I started with a full audit of AI features, existing eval coverage, and the behavioral standards that had actually been defined. From there, I built the evaluation framework around the questions that mattered most: what good looked like, where the model was failing, and which failures were important enough to block or prioritize.
Scoring architecture. I designed a hybrid scoring system that used code-based assertions for deterministic failures like tag presence, format compliance, and length rules, and LLM-as-judge for subjective quality dimensions like coherence, tone, and reasoning. Getting that distinction right is the difference between a reliable eval signal and a score that only looks precise.
Judge calibration. I validated LLM judge outputs against human assessment before trusting them at scale. At BCNY, I built that validation loop directly into the evaluation process, using human review to calibrate judge behavior and reduce the risk of misleading signal.
Representative dataset construction. I build evaluation datasets from actual user behavior, not assumed or hypothetical categories. At BCNY, I analyzed 10,000-plus real queries to create a usage taxonomy that grounded our synthetic test sets in how people actually interacted with the product.
Production monitoring and regression management. I set up automated daily eval monitoring, built a severity classification system that assigned engineering ownership to regressions, and established a practice of testing eval suites against new model versions before release. Visibility without accountability is just reporting; I built both.
Making evals actionable. I treat evals as a product and engineering feedback loop, not a separate reporting function. At BCNY, I owned a weekly insights cadence for product and executive leadership that became the team's primary signal for quality prioritization.
Prompt management. I reviewed, edited, and submitted PRs for system prompt changes as part of the eval cycle, using regression signals to propose targeted fixes to model behavior. Prompt iteration and evaluation were part of the same loop, not separate workstreams.
Eval coverage analysis. I used user feedback signals to identify where existing evals were falling short or missing real failure modes. That closed the loop between what users were actually experiencing and what the evaluation suite was measuring.
Tools: Python, LLM APIs (Anthropic, OpenAI, Google), Braintrust, Retool, BigQuery, Cursor, SQL, GitHub
Product Operations, for me, has always been about finding the friction other people have learned to work around and building the system that removes it. Across Meta, Pinterest, and The Browser Company, that instinct has turned into annotation UIs, LLM pipelines, self-serve reporting tools, Tableau dashboards, and training infrastructure that made teams faster, clearer, and less dependent on heroics.
LLM pipeline design and AI automation. At BCNY, I designed and built a chained LLM pipeline that connected explicit user feedback from BigQuery with implicit feedback from chat transcripts stored in S3. The pipeline ran multi-step summarization and categorization on an hourly cadence, creating systematic visibility into quality patterns and failure modes product and engineering had not previously been able to see, including a routing issue that was ultimately reduced by 80 percent.
Self-serve internal tooling. I build tools that teams can use without needing me in the loop. At Pinterest, I replaced a CURL-based CLI for annotation job submission with a Retool GUI featuring drag-and-drop upload, input validation, and one-click submission, bringing ramp-up time for new PMs to under 15 minutes and reducing launch errors to zero.
Annotation UI development. I designed and built 45-plus annotation UIs in HTML, CSS, and JavaScript across every interaction type our programs needed, including image bounding boxes, multi-select, free text, and chat interfaces. When we migrated vendors, I included a REST API requirement in the RFP that enabled auto-generation of per-row HTML from shared templates, cutting job prep time from hours to minutes.
Reporting and data infrastructure. I've owned reporting end to end, from Tableau dashboards tracking rater alignment for 360-person workforces at Meta to the LLM-analyzed weekly insights cadence I built at BCNY. Clean, decision-ready data does not surface itself. I build the systems that make quality, progress, and risk visible.
PgM enablement and L&D. At Pinterest, I built a dual-purpose training portal when PgMs transferred in from another team with no data labeling background and limited onboarding support. The portal turned scattered institutional knowledge into a reusable resource, reducing dependency on me as the knowledge bottleneck and helping the team ramp faster.
Cross-functional coordination. Operational infrastructure rarely has a natural champion, which means making the case is part of the work. I've navigated vendor budget negotiations, legal and compliance review for trust and safety tooling, and engineering roadmap conversations across Meta, Pinterest, and BCNY. In smaller, more ambiguous environments, I have often been the person translating operational gaps into concrete asks for product, engineering, and leadership.
Tools: Python, HTML/CSS/JS, Retool, Tableau, BigQuery, SQL, Google Sheets + Apps Script, Intellum, Articulate 360, Rise, Notion, Coda
When I'm off the clock, I'm usually playing something or making something. The short version:
Florida-raised, Oakland-based, and more of an extrovert than people expect. If I'm invited somewhere, I'm probably going.
Outside the program and operations tracks I've spent years building, these are three areas I could see myself pivoting to:
| Area | Why I'm drawn to it |
|---|---|
| Product Management | The closest to where I already am. Cross-functional coordination, product sense, tooling, launch readiness. I've lived right next to it for years and can see myself making the move. |
| Design Program Management | The same instinct behind a lot of this site: after years deep in technical work, I want to be closer to the creative side again. Not as a designer, just near it. |
| Product Marketing | Simple, really: I've loved every product I've worked on, and helping promote something I believe in sounds like work I'd be good at and genuinely enjoy. |
If you're recruiting for these roles and open to discussing what that would look like at your company, please reach out.
Let's connect →
Find me on social, or head to the contact page for the form.