Devin Huber

Devin Huber · AI agent builder

I build
AI agents.

that tag the news as it lands

click to add a node · hold to grow it · drag to connecttap to add a node · hold to grow it

scroll
How my agents work

Built for efficiency.
Safeguarded for accuracy.

I build AI agents across the full range, from automating the routine work that eats a team's week to analyzing information, giving feedback, generating ideas and finding connections people miss. Every one has checks that keep it accurate. I went from zero agents a year ago to making 11 in the last month alone, and I'm building toward whatever they can do next.

sample trace · metadata-enrichment
input
dedup
entity
llm tag
guard
AUTO_PUBLISH
EDITOR_REVIEW
SPECIALIST_REVIEW
11agents built in the past month
37.5→16.7%error rate after tuning one agent
3companies given custom spec builds
Agents & workflows

Recent builds

Eight builds, strongest results first. Some are spec builds for hiring teams, some are live work for my current employer, and two are personal projects. Client names are withheld, and figures come from each build's own test runs.

n8n · 11 nodes
Built for
Legal & news information publisher (spec build)
Model
Claude Sonnet via HTTP
Stores
Data Tables (vocabulary)

Metadata Enrichment Agent

For a businessTags incoming news automatically, so editors spend less time sorting articles and more time editing, and wrong tags are stopped before readers ever see them.

Adds controlled-vocabulary subject, industry, geography and company tags to incoming news as it arrives. It decides which articles can publish on their own and which need an editor or a specialist.

  • Near-duplicates are collapsed with 5-word shingles at Jaccard ≥ 0.50 before any tokens are spent.
  • The guard rejects made-up codes, wrong facets and evidence that isn't quoted word for word. Broader terms come from the taxonomy hierarchy, not the model.
  • Tested on a 24-article answer key with planted traps: "shell companies," passing mentions, retired-code bait and a renamed-company alias.
  • The run brief may only cite figures from the scorecard, and the Word export still ships if the brief call fails.
Tuning log, run 42: lowering the industry confidence floor to 0.70 raised straight-through from 20.8% → 45.8%, but silent errors went from 1 of 5 → 3 of 11. I reverted it and re-applied it only after completeness cues were added to catch under-tagging.
Open the evaluation scorecard ↗ Read the full breakdown
37.5→16.7%error rate across 5 test runs
0.85subject confidence floor
100%company-match precision & recall
Claude · scheduled agent
Built for
My own AI job search (personal)
Model
Claude, running unattended
Sources
LinkedIn alerts (Gmail) · Indeed · ZipRecruiter
Schedule
Daily, 10am

Job Board Intelligence Agent

For a businessReplaces hours of daily job-board scrolling with a short list of roles that actually fit. The same setup works for any daily search: sales leads, vendors or candidates.

A scheduled Claude agent that searches three job sources every morning, keeps only roles that are genuinely AI-scoped, never sends the same posting twice, and emails a short, phone-friendly digest.

  • No daily quota. Every role has to clear a written relevance bar, and a day with nothing worth sending produces a two-line email instead of filler.
  • Its own sent emails are its memory: before searching, it reads 21 days of past digests and builds an exclusion list so repeats never go out.
  • LinkedIn has no connector, so it reads LinkedIn alert emails for discovery, checks Indeed for full descriptions, and labels anything it couldn't verify.
  • Built to run with nobody watching: it reloads dropped connectors, never stops to ask a question, and still sends an honest email if every source fails.
Read the full breakdown
180+roles recommended directly, zero repeats
~4,000job postings screened
48hpriority-apply flag
n8n · 6 nodes
Built for
Legal marketing agency (spec build)
Model
Claude Sonnet, batched ×5, 3 retries
Rules
ABA Model Rules 7.1 / 7.4, state variants

Bar-Compliance Content Reviewer

For a businessCatches risky claims in law-firm marketing before it goes live, so reviewers focus on real problems and clients stay clear of advertising-rule violations.

Checks law-firm marketing copy against attorney advertising rules before it goes live. Regex catches the obvious banned terms. The LLM catches what regex can't: implied guarantees, superiority claims and statistics without context.

  • Disclaimers are removed before scanning, because a compliant disclaimer itself contains the word "guarantee."
  • Each client has its own threshold (0.80–0.90) that separates blocking findings from advisory notes. Outcomes are BLOCK, REVIEW_REQUIRED, PASS_WITH_NOTES and PASS.
  • The scorecard shows which layer caught each violation. If the LLM never catches anything regex missed, it isn't earning its cost.
Read the full breakdown
40ground-truth documents
10 · 5clients · state rule sets
7LLM violation codes
Zapier Agent
Built for
Leadership-development institute (current work)
Sources
GA4 (2 properties) · Kajabi
Schedule
Every Monday

Monday Operations Brief

For a businessLeadership gets a weekly read on site traffic and enrollments in their inbox, without opening a dashboard, and only hears about issues worth acting on.

An analytics agent that emails leadership each week with the biggest changes, green and red flags, and anything worth optimizing across site traffic and the course platform.

  • Reads as an informative email, not a to-do list. It offers suggestions and never says "we have to."
  • On quiet weeks the suggestions section is left out entirely, instead of printing "nothing this week."
  • When items do come up, they're ordered by ease and importance together, with no effort or impact labels.
  • Tested against the live account, which surfaced a connector that ignored its tag filter and a report action limited to one date range. Each run has a 10-step budget and a restricted tool list.
Read the full breakdown
Rev 4built from live-account testing
10step budget per run
0filler sections on quiet weeks
n8n · 3 workflows
Built for
Legal marketing agency (spec build)
Schedule
Monthly, per client
Stores
Snapshots + competitor registry

Competitor Watch & Monthly Digest

For a businessGives account teams a 90-second monthly read on what competitors changed, instead of hours spent manually checking websites.

Monitors each client's competitor websites, detects what changed since last month, scores how much each change matters, and writes a digest an account lead can read in 90 seconds.

  • The registry builder finds case-results, attorney, awards and news pages based on what each page is, whatever the site calls it.
  • Change detection strips nav, footer and scripts, compares a djb2 hash, and passes only new sentences to the model.
  • A test fixture trims stored snapshots and backdates them 30 days, so the diff and scoring can be tested without waiting a month.
  • Scoring failures come back as REVIEW_MANUALLY, never dropped. The run reports a signal ratio to catch a threshold set too low.
Read the full breakdown
0.65materiality threshold
5 · 7calibrated bands · categories
90starget read time
Zapier · 2 agents
Built for
Leadership-development institute's podcast
Record
Google Sheets
Output
Scores, decline drafts, monthly digest

Guest Application Flywheel

For a businessCuts hours of manual screening so strong podcast guests aren't lost in the pile, while a person keeps the final say on every decision.

Two agents that score incoming podcast guest applications, draft personal decline emails for weak fits, and send leadership a monthly digest of the pipeline.

  • Safety comes from limiting tools: each agent holds only the actions its job needs.
  • Application text is treated as data. The agents are built to ignore instructions hidden inside applications.
  • Writes happen in a set order, so the sheet never shows a half-processed row. A person can overrule any score.
Read the full breakdown
15point scoring rubric
5scored criteria
2agents with separate tool sets
n8n · form trigger
Built for
AI note-taking product for social care (take-home)
Model
Claude Sonnet, temp 0.1
Companion
Pipeline & quality dashboard

Assessment Template Test Harness

For a businessLets a product team change its AI prompts with confidence: every change is checked against real conversations before it can affect a social worker's records.

Runs any conversation transcript through an assessment template, so prompt changes can be regression-tested on real input instead of eyeballed.

  • Never infers a diagnosis, capacity, risk or eligibility, and never softens a risk someone stated.
  • Contradictions aren't resolved. Both sides are recorded with who said them, and figures are kept exactly as spoken.
  • The companion dashboard runs 15 statistical tests with multiple-comparison correction, recalculated in the browser from raw rows.
Open the companion dashboard ↗ Read the full breakdown
0.1temperature
11self-check items
15corrected stat tests
n8n · 14 nodes
Built for
A family member's house search (personal)
Schedule
Daily 8am · Sunday top 10
Model
None, fully rule-based

House Search Daily & Weekly Digest

For a businessTurns a flood of duplicate listing alerts into a ranked shortlist, so the buyer only looks at homes worth a visit. The same scoring pattern fits leads or inventory.

Turns noisy listing-alert emails from three sites into a star-rated daily digest, then sends a Sunday email with the week's best ten and an interactive board.

  • No LLM. Listing text is scored with weighted condition phrases ("new roof" +3.0, "as-is" −6.0) scaled to 0–12.
  • A 15-term auction and foreclosure filter logs the phrase that caused each listing to be dropped.
  • Dedup runs in three passes: listing ID, then normalized address plus ZIP, then across sources in the same batch.
  • Missing fields score at the midpoint and are flagged on the card, not silently penalized.
Read the full breakdown
126point weighted score
13scoring factors
3-passdeduplication
swipe, drag or use ← →
How I make AI reliable

Different agents. Same discipline.

Tagging, compliance, monitoring, analytics, screening: every build solves a different problem, but each one is held to the same safeguards before it touches real work.

Safeguard 1 · Rules

Rules before reasoning

Regex, dedup and entity matching run first. The model only chooses among grounded candidates. For example, a common-noun check keeps "shell companies" from tagging Shell plc.

Safeguard 2 · Guard

Guards reject, not trust

Model output gets checked before anything moves. Codes that aren't in the vocabulary are rejected as HALLUCINATED_CODE. Quoted evidence that doesn't appear word for word in the source is rejected as UNGROUNDED_EVIDENCE.

Safeguard 3 · Route

Uncertainty goes to a person

Confidence floors are set per client and per category. If an LLM call fails, the item goes to review instead of being dropped.

Safeguard 4 · Output

Quiet by default

Briefs only raise an item when it's clearly needed. Scoring prompts check their own calibration: if most items score above 0.6 you are miscalibrated.

Safeguard 5 · Eval

The agent grades itself

Decision-making builds run against an answer key and report precision, recall, silent error rate, gating false positives and cost per 1,000 items.

HTML builds

Dashboards, reports & build sheets

Hand-built, self-contained pages that go with the agents: evaluation reports, interactive analysis and the specs the agents are built from. Pages that contain client data are shown live on a call.

Metadata Agent Scorecard

Live ↗

Evaluation report for the enrichment agent: results against the editors' answer key, findings from each test run, and a 90-day rollout plan.

  • precision / recall / F1
  • cost per 1,000
  • rollout plan

Notes Pipeline Insights

Live ↗

Interactive dashboard analyzing an AI note-taking product's pipeline and summary quality, filterable by team, template and period.

  • 4 tabs
  • funnel + hours-saved slider
  • 15 stat tests

House Search Board

On request

Interactive map and ranked list of 239 listings, each with a star rating and a score breakdown you can expand.

  • SVG map
  • 11-factor score
  • filters + detail panel

Signal to Pipeline

On request

Diagnosis of a training company's marketing and sales funnel from CRM, LMS and survey data, with seven proposed automations ranked.

  • drop-off analysis
  • cohort revenue
  • phased roadmap

Monday Brief Build Sheet

On request

Build spec for the analytics agent, with a log of fixes found by testing against the live account and a copy-paste instruction block.

  • fix log
  • step budget
  • tool restrictions

Guest Flywheel Build Sheet

On request

Two-agent spec covering the scoring rubric, which tools each agent gets, defenses against instructions hidden in applications, and a test plan.

  • 15-pt rubric
  • injection defense
  • test plan

House Digest Runbook

On request

Setup guide for the listing digest: inbox labels and filters, the parser, the dedup table, and why it uses alert emails instead of scraping.

  • runbook
  • Gmail filters
  • dedup ledger
About

Operations first, then automation

I'm an AI Solutions Analyst in Columbus, Ohio, building and evaluating AI agents for a leadership-development institute.

Most of my work starts with a person doing something by hand every week: pulling numbers, reviewing copy, screening applications. I work out what they check, what they'd never let slide, and when they'd want to be asked. Then I build an agent that follows the same rules, and I measure its accuracy before anyone relies on it.

My degree is in finance (B.S.B.A., The Ohio State University, 2025), so I frame automation in terms of business impact: hours saved, errors caught, cost per item.

Strengths

  • Multi-layer agent design
  • Hallucination & conformance guards
  • Statistical testing
  • Human-in-the-loop routing
  • Evaluation harnesses
  • Accuracy & false-positive measurement
  • Prompt engineering
  • Webhook & API integrations
  • GA4 & KPI reporting
  • Excel modeling
  • CRM workflow design
  • Requirements & process mapping
  • SOPs & enablement

Tools I build with

n8nworkflows · Data Tables · Code nodes ZapierAgents · MCP ClaudeOpus · Sonnet · Haiku Claude Coworkagents · connectors · skills Anthropic APIdirect HTTP JavaScriptin-flow logic HTTP / REST Google Analytics 4 Kajabi Gmail Google Sheets HTML · SVGdashboards
Contact

Want to see one run?

I'm happy to screenshare any of these builds: open the workflow, run the test set, and walk through the scorecard, including the runs that didn't work.