Skip to content

Portfolio · Updated October 3, 2026

I Don't Advise on AI. I Build It.

Agor AI Studio builds AI agents that run one workflow, at a fixed price. This is the work behind that offer: four apps live on the Apple App Store, shipped in seventy-seven days, thirteen live sites, three published books and a research podcast, all built and run with AI. I also built a 39-agent autonomous build pipeline and then retired it, because an adversarial review of my own work found its governance had never once fired. When I advise a client on AI, I have already met the problems they are about to.

Agents

39

Apps

4

Sites live

13

Articles

1016

Episodes

44

Tests passing

1,300+

Archived Flagship

MVAT Studio: an autonomous multi-agent app factory

A framework in which 39 AI agents built, tested and shipped mobile apps, with no human in the loop during pipeline execution. It put real apps on the App Store. It is also archived, and the reason why is the more useful half of the story.

Architecture

Product

5

Strategy, PRDs, Personas, Prioritization, Market Research

Design

5

UX, UI, Design System, Interactions, Accessibility

Engineering

8

Architecture, Frontend, Backend, Security, DevOps, Code Review

Testing

5

Strategy, Unit Tests, Integration Tests, Quality Gate, Auto-Heal

Marketing

5

ASO, Content, Social Media, Ad Ops, Launch Coordination

Analytics

5

Metrics, Behavior, Crashes, Anomalies, Experiments

Finance

4

Revenue, Budget, Forecasting, Spend Alerts

Governance

2

Pipeline Judge, Spec Evolver (mutual oversight)

Tiered model assignment: cost optimization without quality loss

7

Opus

Production code + critical gates

19

Sonnet

Content, analysis, design specs

13

Haiku

Read-only analytics + reporting

Governance innovation

The hard problem in multi-agent systems isn't making agents that work — it's making them fail safely. All governance is versioned JSON with git-based enforcement hooks. Zero infrastructure.

Circuit Breakers

Auto-trip after 3 consecutive failures, pausing agents before errors cascade through the pipeline.

Pipeline Judge

Independent cross-department validator catching goal drift and hallucination propagation at every stage transition.

Mutual Oversight

The spec-evolver and pipeline-judge cannot modify each other. Only the founder can — eliminating self-modification loops.

Confidence Gating

Auto-execute above 0.85, flag for review at 0.65–0.84, escalate below 0.65. No ambiguous thresholds.

Correction-Driven Learning

Founder feedback as the primary learning signal via append-only correction logs — a feedback loop that compounds over phases.

Assumption Registry

Temporal history of system beliefs — tracking what the system believes, when beliefs changed, and why.

10-stage looping pipeline

  1. 01Discovery
  2. 02Strategy
  3. 03Design
  4. 04Engineering
  5. 05Code Review
  6. 06Testing
  7. 07Build/Deploy
  8. 08Marketing
  9. 09Release/Monitor
  10. 10Feedback Loop

A pipeline-judge validated every stage transition. Stage 10 looped back to Stage 1 with a cross-department synthesis report. Six rollout phases (R1–R6) progressively expanded system autonomy. Final state: R6, full collective autonomy, reached before the framework was archived.

Why it was archived

Archived June 2026. A 67-finding adversarial review of my own framework found that agent identity never bound, so every kill-switch, circuit-breaker and rate-limit check had silently passed for the framework's entire operational life. Zero block events were ever recorded. I retired it rather than repair it, and rebuilt the single idea that worked as a machine-wide action-boundary hook, which has since denied more than 150 real actions and logged every one of its bypasses. The apps it shipped stay live.

The Portfolio

Everything I've built

Worked examples, apps, sites, frameworks and books, shipped with AI. Every recommendation I make to a client is something already running in production here.

MVAT Studio

Archived

Framework

The 39-agent autonomous software factory: 8 departments, a 10-stage looping pipeline, and git-based governance that shipped mobile apps end-to-end. The framework itself was archived in June 2026 after an adversarial review and a written retirement postmortem; the apps it shipped stay live, and its governance patterns moved into a machine-wide action-boundary hook.

voice-agent-dotnet

Shipped

System

A C#/.NET 8 port of my phone agent's call loop, built to test the design: G.711 audio, local voice activity detection and barge-in, tool calling, retrieval, caller memory, verbatim recorded disclosures, and call-frequency rules for payment reminders, with 137 tests. I ran the same scripted call through four voice models (xAI, OpenAI Realtime, OpenAI GPT-Live, Gemini Live); xAI replied fastest. Testing it against live models found three bugs in my production line, fixed the same day.

MVAT Mirror

Live

App

Zero-question personality profiling from real-world behavioral signal — no quizzes, no self-reporting. Live on the Apple App Store with resumable full-history import and credit-based pass-through pricing.

Coqui Chorus

Live

App

iOS nature-soundscape app built on bioacoustic synthesis models — synthesizing sleep/wellness audio from scientific data, not loops. Live on the Apple App Store.

Showing 12 of 50 projects

Shipped Products

Apps built by the pipeline

MVAT Focus

Focus Timer · iOS

Pomodoro-style focus timer with free and Pro tiers. Live on the Apple App Store, with Pro sold as an in-app purchase and a Stripe web tier alongside it.

  • Expo SDK 52, TypeScript strict, Firebase
  • Apple Sign-In + Google OAuth
  • Pro: $0.99 lifetime in-app purchase, plus a Stripe monthly or annual web tier
  • 287 passing tests wired to CI, 0 type errors
  • App Store: live

MVAT Mirror

Personality Profiling · iOS

Zero-question personality profiling from real-world behavioral data. No quizzes. No self-reporting. Just signal from music, browsing, and purchase patterns.

  • Expo SDK 55, TypeScript strict, Zustand
  • Live on the Apple App Store (1.0.2)
  • Resumable full-history import with rate-limit cursors
  • Credit-based pass-through pricing for imports
  • Consumable import credit packs: $1.99 Starter, $4.99 Full History
  • 572 passing tests wired to CI, 0 type errors

Infrastructure & Automation

Build over buy

Systematically replacing subscription SaaS with self-hosted solutions. Full ownership, zero ongoing cost, better observability.

Self-Hosted Link Tracking

Replaced a $30/mo SaaS link tracker with a 15-line route that 301-redirects with UTM parameters. Same functionality, zero ongoing cost.

Dub.co → self-hosted

Auto Social Posting Pipeline

Blog posts auto-syndicate to X, Facebook, and LinkedIn on deploy. Detects new posts via blob-stored manifest, generates tracked links, prevents duplicate posts. X now falls back to the X API when the browser post fails.

Zero-touch publishing

Pre-Call Brief & Prospect Research

Every booking triggers a claude -p recon brief emailed before the call, now including a fingerprint of the prospect's company tech stack; a twice-weekly outbound agent scores inbound-fit prospects from news and funding signals.

Sales prep, automated

Multi-Site Orchestration

Multiple live websites updated simultaneously using parallel sub-agents. Cross-repo changes, deploys, and live verification in a single session.

Hours → minutes

Browser-as-API Automation

When platforms lack APIs, automated via Playwright — treating the browser as a programmable interface. Gmail aliases, store configs, OAuth setup.

35 Gmail aliases automated

SEO Flywheel & Instant Indexing

A weekly Search-Console-driven loop feeds real search demand into the blog generator, now with per-property keyword queues and a suppression list that retires clusters no longer worth chasing. A daily IndexNow sync pushes new URLs to Bing, Yandex, and DuckDuckGo, and a weekly written audit reads the numbers before they frame a decision — which is how a 917-impression jump turned out to be a subdomain counted inside its own parent.

Compounding organic reach

Revenue Loop & Merchant Watch

A weekly loop on modelstack.digital measures the funnel end to end before proposing work: Search Console rankings, server-side GA4 purchases with real attribution and drift gates, and a year of Stripe history. A Merchant Center watchdog checks the actual feed and waits a cycle before calling a product lost, after a shrinking catalog set off a false alarm.

A ranking problem, not a checkout problem

Agents on Real Phone Numbers

Customer agents answer provisioned phone numbers with per-agent voice-minute metering, per-caller and global rate limits, voicemail-aware answering, and bookings that survive a caller hanging up mid-confirmation.

PSTN, not just web chat

Fleet Health Monitoring

Every scheduled task, repo, and the knowledge brain report into one health file surfaced on a local dashboard, alongside a weekly CI audit across all 98 repos. Built after silent failures ran for weeks unnoticed: a 20-run workflow failure cluster nobody saw for a day and a half, a stray carriage return that exited 255 every run, and a launcher that hung waiting on a package registry.

Silent failures surfaced

Technical Depth

Stack

AI / ML

  • Claude API (Opus / Sonnet / Haiku)
  • OpenAI
  • Google Gemini
  • ElevenLabs
  • HeyGen
  • xAI Realtime + Grok TTS
  • Gemini Live
  • OpenAI Realtime / GPT-Live
  • Ollama (local models)

Mobile

  • Expo / React Native
  • EAS Build & Submit
  • App Store Connect
  • react-native-iap

Cloud

  • Firebase (Firestore, Functions, Auth)
  • Netlify (Functions, Blobs, Deploy Hooks)
  • Supabase (multi-tenant Postgres)
  • Cloudflare (Workers, D1, R2)
  • Twilio (Voice, SMS, A2P 10DLC)
  • GitHub Actions CI/CD
  • Stripe (Payments, Subscriptions, Webhooks)

Languages & Frameworks

  • TypeScript (primary)
  • Python
  • Next.js / React
  • Astro
  • Node.js
  • C# / .NET 8 (ASP.NET Core)
  • Bash / Shell scripting

Auth & Security

  • Apple Sign-In
  • Google OAuth
  • Firebase Auth
  • OAuth 2.0 / PKCE
  • OWASP security patterns

Data & DevOps

  • Postgres + pgvector
  • Firestore / NoSQL
  • Playwright automation
  • MCP servers
  • Git-based governance
  • OTA updates (Expo)

Thought Leadership

1016 published articles

Writing at the intersection of AI strategy, autonomous systems, and organizational design — across agor.me, modelstack.digital, and scored.tools. A new essay ships nearly every day.

Read all articles

Agor AI Podcast

44 podcast episodes

Weekly deep-dives into the latest AI research papers — what they mean for strategy, automation, and the future of work. Scripted from the week's papers and produced with AI voices, including a clone of my own.

13:49

AI Papers Weekly: Cheaper Agents, Safer Code, Tougher Web Defenses

Three new studies tackle the practical limits of AI agents: turning costly model skills into cheap reusable tools, keeping humans in control of AI-written software, and training web agents to resist hijacking. Each offers lessons on cost, governance and security.

16:01

AI Papers Weekly: Closing the AI Say-Do Gap & Managing Risk

This week, we explore the hidden risks in deploying autonomous AI. We uncover the 'say-do' gap where agents fail to execute their plans, expose the illusion of AI explainability, and reveal a breakthrough in training risk-averse models. Essential listening for leaders scaling AI safely.

10:40

AI Papers Weekly: The Vocabulary, The Stopwatch, and The Poisoned Tool

A randomized trial clocks Figma Make cutting design time ~20%, with PMs gaining most. A governance paper argues psychological words like 'trust' and 'memory' misgovern agents. And a black-box attack hijacks MCP agents 93.6% of the time via poisoned tool metadata.

16:28

AI Papers Weekly: Governing Agents, Shadow IT, and Data Strategy

This week, we explore the governance and security risks of deploying multi-agent systems across enterprise boundaries. We also tackle the growing "shadow IT" problem of downloadable agent skills and settle the debate on whether to use RAG, prompt caching, or fine-tuning for proprietary data.

13:50

AI Papers Weekly: When Agents Cave, Catch Up, and Out-Diagnose Doctors

This week: LLMs abandon correct answers under sustained user pushback, a new recipe lets smaller models match frontier performance at a fraction of the cost, and a clinical AI beats physicians 82% to 57% on primary-care diagnosis.

14:17

AI Papers Weekly: When Agents Hide, Cheat, and Invent Their Own Language

Three papers cut through the AI agent hype: a Nobel-caliber framework for contracting with agents that can lie about their capabilities, an audit exposing how guardrail 'welfare gains' were measurement artifacts, and evidence that multi-agent LLMs spontaneously evolve languages humans can't read.

The studio

How an engagement runs

One workflow at a time, at a fixed price. Every recommendation is something already built and running here.

01

AI Agent Architecture Review

$2,000

The architecture (the systems the agent connects to, what it handles on its own, and where a person signs off), a build plan for that workflow, a cost and risk assessment, and a fixed quote for the build. Credited in full against the build if you go ahead within 30 days.

02

30-day Agent Build

$15,000 to $25,000

Fixed price, quoted in the review. It delivers one production workflow, live in 30 days: an inbox agent, an AI phone receptionist, reporting automation or lead handling.

03

Run & Improve retainer

$3,000 to $6,000 a month

Monitoring of the agents we built, iteration as your work changes, and new workflows one at a time. Scoped with you once the build is live.

Let's build something real

Whether you need AI strategy, multi-agent architecture, or hands-on implementation — I've already done it. Let's talk about your challenge.