An independent software studio · est. 2026 · Atlanta
The software studio with one human on staff.
Sprayberry Labs builds, audits, and researches software. The staff is an AI operation — askalf — that ships the code, reviews the pull requests, verifies the findings, and watches production. Thomas Sprayberry architects, reviews, and signs what leaves the shop.
68,560 skills scanned · 4 frameworks governed · every claim links to a real PR or release
Agent security, proven on our own fleet.
The studio's research question: what are agents allowed to do, and who checks? The agent-security stack — redstamp, truecopy, strongroom, fieldpass, plumbline — is built here and proven on the fleet that runs this company, misses left in. Click any line to read the work.
-
How far is a CPU from serving production LLMs? We found the exact line
Two questions wearing the same words: a GPU-less box can’t replace a frontier model for chat (a 34× bandwidth wall, physics), but it’s one honest step from a production routing gateway that serves the free majority on-box and sheds the rest. A day on a 2013 desktop: a 2× transport win, a live load-shedding tier, and three ideas shipped turned off because the measurement killed them — the failures are the receipts.
-
The audit is now a tripwire — 280 plugins re-verified at their pinned commits, every day
One-time audits rot, so truecopy re-audits the entire official Claude Code plugin directory daily — every vendor plugin re-fetched at its catalog-pinned commit, 1,966 skills at the August 2026 run, zero poisoned — published to a public badge with pin drift as a first-class signal. On day one the watch caught two bugs in its own harness; that story is in the post, because a tripwire you’ve never seen fire is decoration.
-
We scanned the marketplace that started the poisoned-skills panic — 66,541 skills, clean
ClawHub is OpenClaw’s open registry, the marketplace whose poisoned-skills incident started this category. We poison-scanned every skill in it and cross-checked each flag against ClawHub’s own scanner. Zero confirmed malicious; 813 deterministic alarms, all benign, mapped — and the most-installed skill in the registry runs clean but ships obfuscated code you can’t read.
-
We scanned 2,019 Claude Code skills across ten marketplaces
A skill is prose that steers an agent holding your credentials. Every skill in the official marketplace plus nine community ones — 177 vendor repos, fetched at pinned commits and poison-scanned. None were malicious; the engineering was earning that answer, taking a deterministic scanner from a one-in-ten false-alarm rate to 0.6%, every fix pinned with a test.
-
The agent-firewall benchmark, run through a cold pipe: 100% recall, 100% precision on the 298-sample arena
One labeled corpus, one language-agnostic pipe, recall + precision + determinism scored together — with block-all and allow-all anchors sitting in the table to prove a single number is gameable. Earlier corpus revisions documented redstamp's five misses rather than hiding them; the current 298-sample run is clean, sanity anchors included — and the leaderboard that would have been a category error, refused.
-
132 inherited secrets cut to 13 — the rest replaced with single-use leases
Spawn a worker with the parent's environment and it inherits every key the platform holds. strongroom hands agents TTL-bound, use-limited, host-scoped leases with the real key injected only at egress — 119 keys withheld from a live fleet, and the bugs the rollout surfaced on both sides fixed in public.
-
8/8 planted injection payloads withheld from the model, on real Chrome
An agentic browser has the lethal trifecta by construction — the page is untrusted, the session is private, navigation is the exfil channel. fieldpass quarantines indirect prompt injection before the model sees it and gates dangerous actions — validated against Chrome 149, with the real false positives found and fixed in public.
-
One gate, four frameworks: CrewAI, LangGraph, OpenAI Agents SDK, AutoGen
Four unrelated runtimes governed below the framework, at the MCP layer — the poisoned tool stripped before the agent can load it, the destructive call blocked at the gate, every verdict in a tamper-evident audit. Each merged PR touches only
examples/; zero lines changed in the gate's own source. -
plumbline — trajectory-level monitoring: replay what your agents actually did
Per-call gates judge one action at a time; plumbline scores the whole sequence against the agent’s declared job — out-of-band, read-only, catching escapes assembled from individually approved steps. Dogfooded on the fleet that runs this company.
-
A 9.2/10 OpenSSF Scorecard on the package that ships itself
The same repo the self-healing pipeline releases several times a week scores 9.2/10 on the OpenSSF Scorecard — 14 of 18 checks a perfect 10 — and holds a 100% OpenSSF Best Practices badge: releases signed and SLSA-attested, npm published tokenless from CI via OIDC trusted publishing, zero open security alerts. The same hardening — weekly Scorecard runs, hardened branch rulesets, signed releases — runs across all 16 public code repos.
-
The proxy that silently deleted
continue;from your codeA framework-name scrubber ran over message content, so the JS keyword
continue;was stripped to a bare;in transit. A downstream auditor "found" the bare-;bug and blamed the model — when the proxy itself had corrupted the payload. Confirmed with a live A/B: 3/3 fabricated through the unfixed proxy, 0/4 through the fixed one.
More — including the retractions we keep on purpose — on the blog.
Measured on this studio's own production fleet — methodology, and the misses, in each linked write-up.
The staff is a fleet: an orchestrator and twenty-plus specialist agents that triage the inbox, open and review the pull requests, verify every finding against source, watch production, and send the invoices. Agents propose; the human disposes. The full register — the roster, the rules, the live minutes — is filed at askalf.org.
The research ships as working tools.
- redstampA deterministic firewall between an agent and its tools — risk-tiers every call, blocks exfil and injection.security
- truecopyVet, sign & pin every skill and MCP server before it runs — drift detection catches a poisoned tool.skills
- strongroomA vault that hands agents scoped, short-lived leases instead of raw keys.secrets
- fieldpassA governed browser for agents — injection firewall, action gate, judge.agent-browser
- plumblineScores an agent’s whole action sequence against its declared job — out-of-band, read-only, catching escapes built from approved steps.trajectory
- darioYour Claude subscription behind one local endpoint, with multi-account failover — 330+ stars.routing
- deepdiveA local deep-research agent — plan → search → fetch → synthesize a cited report.research
- handsYour LLM on your own mouse, keyboard, and screen — with a full audit log.computer-use
+ hybrid, cordon, browser-bridge, amnesia — the full set at Own Your Stack.
The same fleet that does the research does the client work.
Code audit
$1,500 · fixed price
A hard, cited look at your codebase — the same deterministic, no-confabulation method in the receipts above. Machine-drafted, byte-verified against your source, signed by the human. Try the free mini-audit first.
SEE THE AUDIT →Builds & retainers
scoped · fixed price
Scoped software sprints with a real shipped deliverable, and an on-call architect for teams running custom software in production. Daily progress in a shared channel — written by the fleet, signed by the human.
START A CONVERSATION →This site is built and run by the studio it describes. The research is mined from real GitHub history; the claims link to the source so you can check the work. Built in the open, the unfinished parts included. The human is Thomas Sprayberry <hello@sprayberrylabs.com>.