Real AI or a wrapper: how to tell during diligence
A wrapper is not automatically a bad business, but you should price it as one. How we separate real AI from a thin layer during due diligence.
Moqa Stories
Practical articles on AI systems, technical audits, and due diligence, from the team that ships them. In English and in French.
A wrapper is not automatically a bad business, but you should price it as one. How we separate real AI from a thin layer during due diligence.
Seats are 19 to 40 dollars, tokens can be ten times that. The real 2026 budget for coding agents at team scale, and an ROI frame that survives scrutiny.
Claude Code, Cursor and Copilot are different tools, not rivals. What each does best, what they cost in 2026, and how we choose per team and policy.
How to reduce LLM costs in production: model routing, caching, context discipline and batching, and why unit cost must be designed before you scale.
Running multiple coding agents in parallel: worktrees, task isolation and review queues, what the pattern demands from your process, where it breaks.
How to verify an AI startup’s performance claims inside a deal window: eval sets, held-out tests, production logs, and reading the demo choreography.
Most AI proofs of concept never ship because they prove the demo, not the workflow. What production requires, and how to scope a POC that can graduate.
The AI coding adoption gap is not a talent problem. It is workflow knowledge stuck in two heads, and it closes with skills, conventions and champions.
AI writes code faster than your team reviews it. Why rubber-stamping and reviewer fatigue follow, and the redesign that gets throughput back safely.
Which process to automate with AI first: four criteria that predict success, why impressive demos make bad first projects, and a scoring shortcut.
A failing test is the only feedback a coding agent reliably respects. How TDD makes delegation safe, who writes the test, and where the leash ends.
Copilot acceptance rate tells you code was inserted, not that delivery improved. The metrics that track real AI impact: cycle time, review, defects.
The baseline metrics to capture before AI adoption: cycle time, review load and defect rate per team, from tools you already run, in a single week.
What LLM observability looks like in production: tracing, cost per request, quality signals, drift and alerting. Why we refuse to ship AI without it.
How to measure AI coding productivity with metrics that hold up: cycle time, review load and defect rate against a baseline. And what to ignore.
Senior engineers skeptical of AI coding are reading the risks correctly. Why delegation multiplies their kind of work, and how their standards scale.
Why test-writing is the ideal first task to hand a coding agent: verifiable output, bounded blast radius, real payoff. And how to keep the tests honest.
What RAG is in plain business terms: why a model needs your documents, what retrieval fixes, what it does not, and when it is the right architecture.
The CI guardrails that make AI-generated code safe to merge: test gates, static analysis, security scans and review rules that cannot be skipped.
Free-form prompting produces free-form quality. Why teams need shared conventions for framing agent tasks, versioned in the repo like linter rules.
How production AI systems control wrong answers: retrieval with citations, constrained outputs, review where stakes are high, and refusal by design.
Redesigning the development workflow around coding agents: tickets as executable tasks, review tiers, CI guardrails, and the touchpoints that stay human.
Veracode found 45 percent of AI-generated code introduces an OWASP Top 10 flaw, and bigger models do not help. The pipeline that actually does.
The six-area checklist we use to audit AI companies for funds: production reality, error handling, evals, unit economics, data rights, team risk.
The rollout playbook for AI coding tools: a baseline week, a pilot squad on real backlog, measured deltas, then champions. Licenses alone change nothing.
How we assess OpenAI dependency risk in AI deals: the four real exposures, what a working abstraction layer looks like, and when dependency is rational.
Where RAG assistants actually go wrong: chunking, retrieval, stale documents, prompt assembly and missing evals, in the order we check them.
How to test an AI startup’s gross margins in diligence: the full cost of a request, sensitivity to model prices, and what to ask for in the data room.
Bans produce shadow AI. The one-page AI coding policy that works: approved tools, data boundaries, review rules that scale with risk, honest attribution.
How to evaluate AI agent startups in diligence: task completion rates, blast radius of errors, guardrails, cost per completed task, and liability.
Chatting with a coding agent caps you at your own reading speed. Delegation is the jump most teams have not made, and it is a learnable skill.
How to give an AI agent access to your systems safely: blast radius analysis, least privilege, action allowlists, approval gates and audit trails.
What an eval set is, why testing LLM outputs by hand fails at the second prompt change, and how much evaluation a production business system needs.
MIT found 95 percent of GenAI pilots show no return, RAND puts AI failure at twice the IT rate. The causes are organizational and the survivors share a pattern.
Building an internal AI assistant on company knowledge: permissions, citations, freshness and adoption, and why the SharePoint chatbot disappoints.
How to measure AI ROI defensibly: baseline first, the metrics that hold up, the costs everyone forgets, and why most reported ROI numbers are soft.
How we evaluate AI moat claims in due diligence: the data loops, workflow depth and switching costs that survive model commoditization, and what does not.
What AI coding assistants do with your code: training and retention policies by vendor and tier, plus the controls to set before rolling out agents.
Agent skills are versioned instructions a coding agent loads for recurring tasks. What they are, which ones pay off first, and how to build a library.
Data readiness for AI is a per-workflow question, not a company-wide prerequisite. What each use case really needs, and what can stay messy.
Why one AI champion per engineering squad beats a central AI team: how to pick them, what they own, and the protected time budget that makes it stick.
How to scope an AI project: one workflow, one metric, one owner, and the smallest system that proves value on real data. With the anti-patterns to avoid.