The model writes code
FunCodeGenerator: sketch a wireframe, and a vision model returns a working prototype you can annotate and iterate.
Systems, not demos.
Agents, MCP servers and gateways, tool governance, context budgets, evaluation that can fail — designed as systems, not demos.
Word 01/03363 unit tests guarding a world that is generated, not authored
363 — unit tests guarding a world that is generated, not authored.
Ibuildsystemsattheintersectionofsoftwarearchitecture,cloudinfrastructureandartificialintelligence—andinvestigatewhatcomesnext.
Enterprise AI asset gateway
2026Own product · in progressPhase 0 complete · Phases 1–8 designed
MCP gateways centralize tools but flood the harness’s context: too many tools, verbose schemas, giant results, flows that need many skills.
A Rust data plane is the only MCP ingress and only enforces; a Python optimizer compiles per-identity context bundles; a control plane on Postgres + pgvector owns policy and audit. Nine contracts fix the seams.
57 tests written first and green, an end-to-end harness with 24 asserts and a negative control that must fail.
Put the invariant in one place.
How the thinking runs — shown by quoting the documents it produced. Spanish originals stay next to the translation.
Name the real constraint before choosing a stack.
Quipu didn’t start with Rust. It started with a diagnosis: harnesses run out of context in four distinct ways. Each got its own defence, and the architecture followed from the defences.
“Governance is context optimization.”
La gobernanza ES la optimización de contexto.
$ visible_tools = policy(role, workspace, assets)Quipu · docs/architecture.md
Fix the contracts before the code.
Nine contracts define how the gateway, optimizer, control plane and agents talk — before most of them exist. The seams are agreed first, so the parts can be built in parallel, by people or by agents.
“Contracts are stable: changing one is a breaking change.”
Contratos (estables — cambiarlos es un breaking change).
Quipu · MASTER-PLAN §6.2
Write decisions down, with the reason and the owner.
Some choices are closed on purpose. Writing them down stops them from being re-litigated in every session — including sessions with AI agents that weren’t there when they were made.
“Fixed decisions — don’t re-litigate without the owner.”
Decisiones fijadas (no re-litigar sin el usuario).
Quipu · MASTER-PLAN §4
Make the test able to fail.
The end-to-end harness ships with a negative control: run it with the wrong allow-list and it has to go red. If it stays green, the harness is what’s broken.
“A harness you have only ever seen pass has proven nothing.”
Un harness que solo se ha visto pasar no ha demostrado nada.
$ CURATED=add,get_time,echo bash dev/smoke-e2e.sh # must FAILQuipu · dev/README.md
State the honest limit.
Deployment artifacts that were never executed are marked as verified by inspection only. Targets a phase can’t meet by construction are exempted — and the plan names the mechanism that will.
“Verified only by inspection: they do not count as tested.”
Verificados solo por inspección: no cuentan como probados.
Quipu · MASTER-PLAN §13
Task“Why did last night’s deploy fail? Open an issue with the root cause.”
96 tools with full schemas + 6 skills loaded in full — vs the 9 tools this role may use, as short signatures.
Can a procedurally generated building feel wrong — instead of just random, or just decayed?
Wrongness is semantic. Noise reads as decay; a correct memory of a place with a few precise errors reads as dread.
Twelve memory errors — a clock without hands, an EXIT sign on a blank wall, carpet climbing the wall — at most three per chunk. A reachability test guarantees no error ever blocks the route.
Does a procedurally generated reverb actually sound like the room it claims to be?
If each room’s impulse response is generated per material, its RT60 per band should measure within spec.
“Dark” materials measured bright. The tail is now three independent noise bands through Linkwitz–Riley filters, and per-band RT60 measures within a few percent of spec.
Should the Phase 0 walking skeleton already meet the data plane’s latency SLOs?
A skeleton should be held to the final targets from day one, so regressions are caught early.
A p50 ≤ 1 ms budget containing a network call is impossible arithmetic. Phase 0 is exempt from two SLO rows, and the plan names the snapshot that restores them.
How should a client’s 3D house configuration be saved?
The obvious design: export the model — geometry, textures — and store the file.
All geometry is procedural. A saved model is ~300 bytes of JSON, validated on read; the price is recomputed, never trusted from storage.
Is a deterministic, identity-projected tool surface better than letting the agent search for tools?
Yes: a surface computed from identity costs fewer tokens and fails less than a search the agent has to drive.
Pending — the Context Ledger (Phase 3) will count tokens saved per day, team and asset.
Earlier investigations, as the earlier portfolio records them.
FunCodeGenerator: sketch a wireframe, and a vision model returns a working prototype you can annotate and iterate.
Quasar: a local Llama 3 8B writes Terraform for AWS, treats apply errors as observations and loops until it works.
KODA: multi-agent DevOps workflows with LangGraph, A2A and MCP — and research on cache-augmented generation.
Quipu: the gateway that decides what a harness gets to see — governance and context optimization as one mechanism.
NOCLIP: a game built by parallel agent workstreams under contracts, verifiers and independent research passes.
0 1 2 3 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 +-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+-+ | OBJECTIVE | +---------------------------------------------------------------+ | FILES TO CREATE / TOUCH | +-------------------------------+-------------------------------+ | CONTRACTS CONSUMED | CONTRACTS PRODUCED | +-------------------------------+-------------------------------+ | TDD: test first · watch it fail · minimal code · green | +-------------------------------+-------------------------------+ | VERIFICATION COMMAND | ACCEPTANCE CRITERIA | +-------------------------------+-------------------------------+
NOCLIP was built by five parallel workstreams — level generation, the camcorder post pipeline, audio, entities and UI — each against its own dev harness. Verifier passes flagged: log soft-lock · seed links · Lessee ambient · pacing · line of sight.
Brainstorm → an approved spec → a phase plan of bite-sized TDD tasks. The plan is the entry point for any new session, human or agent.
Each task goes to a new subagent with only its packet, then a two-stage review before the next one starts.
“Done” means the commands ran and the evidence is shown — tests, lint, the deliverable actually running. Never a claim.
Playwright with fixed seeds and scripted routes: screenshots, FPS logs and a debug hook let agents verify what they built.
For open design problems, independent research passes from different agents against the same build — then a synthesis into contracts.
Lessons go into NOTES.md — one per entry, updated instead of duplicated — and CLAUDE.md / AGENTS.md live in every repository.
I’m looking for teams building AI systems that have to work in production — where architecture, cloud and AI can’t be pulled apart. Based in Quito (UTC−5), working in Spanish or English.
Start a conversation→Agents, MCP servers and gateways, tool governance, context budgets, evaluation that can fail — designed as systems, not demos.
Event-driven and serverless, containers on ECS or EKS, the right store for each access pattern — with the trade-offs written down.
Contracts, SLOs, decision records, honest limits and a verification path — and then building it, phase by phase.
Skills, subagent workflows, verifiers and memory, so what agents produce is something your team can review and trust.