Yenson Umaña · AI architecture & technical strategy for startups

AI architecture for startups.

From AI ambition to production clarity.

I help founders and CTOs decide what deserves to be built, design the production architecture behind it, and give their team a path they can execute.

Senior AI Solution Architect supporting Microsoft for Startups globally via Accenture.

The startup reality

Innovation moves fast. Architecture decisions compound faster.

Startups rarely lack AI ideas. They lack certainty about which decisions will compound. Models change, customer assumptions move, teams grow, and prototype shortcuts quietly become infrastructure.

So the challenge is not choosing the newest technology. It is deciding what matters now, what can wait, what must be validated, what cannot fail, and what should stay flexible while the answer is still unknown.

01

The decision, before the technology

Define the outcome, the constraints, the cost of failure, and what must be true before the initiative deserves engineering investment.

02

Tradeoffs made explicit

Model, data, retrieval, evaluation, security, latency, and human review decided deliberately — with the reasoning written down and reversible where it should be.

03

A path the team can execute

Architecture decision records, sequencing, and acceptance gates your engineers can act on without translating founder intent through three layers.

Where I come in

I join when the team has momentum, but no map.

Early companies rarely arrive with clean requirements. They arrive with customer signals, investor pressure, a changing product, a working demo, and a dozen decisions nobody has time to untangle.

I work alongside the founder and engineers until the next production decision is clear, documented, and owned by the team.

01
Many promising ideas
One prioritized AI bet
02
A prototype that works once
An evaluated production path
03
Model and platform noise
Explicit architecture decisions
04
Context held by the founder
A roadmap the team can execute

Ways to work together

Structured ways to get senior judgment into a decision that is already moving.

These are formats, not packages. The work is the same underneath: frame the decision, make the tradeoffs explicit, and leave the team with a path they can execute. Scope follows a discovery call.

01

Signature

2-4 weeks

AI Architecture Sprint

Turn one consequential AI product or system decision into a production-ready architecture and a delivery path the team owns.

Best for: Teams with something important enough that the architecture cannot be improvised: an agent, a retrieval system, or an AI-native capability moving toward production without a senior architect in-house.

  • Target architecture and architecture decision records
  • System boundaries, guardrails, and human review paths
  • Evaluation, observability, latency, and cost plan
  • Phased delivery sequence and technical handoff
Explore fit for AI Architecture Sprint

02

1-2 weeks

AI Direction Sprint

Answer which AI opportunity actually deserves to become a product initiative — and why the others can wait.

Best for: Founders and CTOs with strong market insight, several plausible AI directions, and no shared basis for choosing what deserves to ship first.

  • Opportunities scored by value, feasibility, uncertainty, and risk
  • Readiness and capability gap assessment
  • One prioritized bet with a defensible rationale
  • 90-day sequence with owners and decision gates
Explore fit for AI Direction Sprint

03

Ongoing, usually after a sprint

Fractional AI Architect

Senior architecture judgment embedded beside the founder and engineering team, before the company needs or can justify the role full time.

Best for: Startups making a steady stream of consequential technical decisions — model and vendor choices, production reviews, roadmap tradeoffs — with no one senior enough to hold the architecture.

  • Founder and architecture working sessions
  • Model, vendor, and platform decisions
  • Architecture and production reviews
  • Technical risk, evaluation standards, and team unblockers
Explore fit for Fractional AI Architect

04

3-6 weeks

AI Engineering Enablement

Turn architecture decisions into shared engineering practice, so the same lessons stop being relearned team by team.

Best for: Scale-ups already running AI in production, but with inconsistent evaluation, unclear ownership, and no shared review standard.

  • Shared architecture principles and review standards
  • Evaluation and production workflow patterns
  • Ownership model and right-sized governance
  • Office hours and adoption measures
Explore fit for AI Engineering Enablement

A good fit when

  • There is real momentum and a customer or market signal
  • AI is strategically relevant, not decorative
  • Several plausible technical directions are open
  • Architecture decisions are becoming consequential
  • The team is capable but has no senior AI architecture leadership
  • The company moves too fast for a traditional consulting engagement

Probably not me

  • A chatbot or a basic automation with no product ownership behind it
  • Generic AI training disconnected from an operating change
  • Inexpensive development capacity
  • Implementation resources rather than technical judgment

One active architecture or direction sprint at a time, plus a small number of advisory relationships.

Startup work

I know what the room feels like before the path is obvious.

Startup architecture means making consequential decisions with an incomplete map. Each case below is written as the decision it actually was, not the outcome it became.

Anonymized startup engagement

An urgent founder objective, turned into a target architecture the team could execute safely.

full migration
7 days
visible downtime
0 min
team handoff
Full

Startup migration · Production path

7 days
  1. 01

    Urgent founder objective

  2. 02

    Target architecture

  3. 03

    Agent-assisted execution

  4. 04

    Acceptance gates

Ambiguous migration

Team-owned system

Context
A US-based Y Combinator startup needed its production stack off Vercel and Google Cloud and onto Azure, on a timeline set by the business rather than by engineering comfort.
Decision
Whether to move incrementally and carry two platforms for months, or define a single target architecture with acceptance gates and cut over once.
Constraints
No customer-visible downtime, a small team that still had to ship product, and no appetite for a migration that quietly became a quarter-long project.
Architecture
A defined target architecture, explicit acceptance gates per subsystem, and an agent-assisted execution path the engineering team ran themselves.
Tradeoffs
Chose one decisive cutover over a long dual-run: higher preparation cost, far lower carrying cost and ambiguity. Deferred optimization work that would have widened scope without reducing risk.
Result
The full stack moved in seven days with zero customer-visible downtime.
Ownership
A runbook and an operating approach the team kept using after the engagement ended.

Named client · Amplification Of Potential

A live-event AI system designed around the conversations it must never create.

p95 latency target
< 2.5s
inference budget / event
< $1
PII in model payload
Zero

AOP Beacon · Decision flow

Guardrailed

Phase objective

01

Guardrail policy

02

Few-shot + fallback

03

Private client render

04

Attendee selections

Safe conversation prompt

Context
AOP wanted AI-generated conversation prompts at live events, where the audience is present, the moment is unrepeatable, and a bad output is visible to everyone in the room.
Decision
How much to delegate to the model at all — and where a deterministic system had to remain in control of what an attendee could see.
Constraints
Under 2.5 seconds at the table, no attendee names in the model payload, no sensitive topics during the early event phase, and an inference budget that had to stay under a dollar per event.
Architecture
A phased experience model, a client-approved few-shot bank, explicit topic guardrails, a privacy boundary that keeps identity out of the payload by construction, and a deterministic fallback path.
Tradeoffs
Accepted a narrower generative range in exchange for bounded failure. Chose approved examples over open generation, and a deterministic fallback over a retry that could miss the moment.
Result
A reusable architecture for six active tables and up to 40 attendees, within the latency and cost envelope.
Ownership
Guardrail policy and prompt bank the client can extend without re-engineering the system.

Founder / operator proof

I also live with the decisions after launch.

Founder and operator

Presencia Loyalty

Built and operate a live wallet-based loyalty SaaS — a product whose architecture decisions I still live with every week.

Connected product experience

Junior Rodríguez × Presencia

Designed an NFC-enabled painting experience that joins a physical object, its story, and a shareable digital journey.

How I work

Founder context, architecture, and execution stay in the same room.

Every engagement produces decisions the technical team can execute without losing the founder's product context. Tools and models are selected after the operating problem is clear.

  1. 01

    Frame the decision

    Before choosing a model or a platform: define the business outcome, the constraints, the cost of failure, and what must be true for the initiative to deserve investment.

  2. 02

    Make tradeoffs explicit

    Compare viable approaches across quality, latency, cost, security, maintainability, and what the team can realistically operate.

  3. 03

    De-risk with evidence

    Use focused prototypes, evaluations, and failure-mode reviews where they resolve an important unknown — not as theater.

  4. 04

    Transfer the capability

    Leave the decisions, the reasoning, and the standards with your team, so the work compounds after the engagement ends.

Model and vendor independence

Architecture preserves options while uncertainty is high

Human review where consequences demand it

Evaluation before automation confidence

Ownership transferred to your team

Point of view

Decision-making under technical uncertainty.

The hard part is rarely finding another AI idea. It is deciding which one deserves to become a system. These are the questions I keep returning to with founders and engineering leaders.

01

Architecture decisions

When an agent should actually be a deterministic workflow. Choosing between retrieval, tools, and fine-tuning. Which decisions become expensive to reverse.

02

Production AI

What separates a demo from a system you can be responsible for: evaluation, latency and cost budgets, observability, and failure modes named before launch.

03

Founder and CTO decisions

What not to build. How to prioritize AI opportunities. When to hire an AI engineer versus an architect. What belongs in the first 90 days.

04

AI economics

Cost per successful task, model routing, inference economics, build versus buy, and what a tolerable failure actually costs the business.

05

Architecture teardowns

Realistic technical situations worked through end to end — the constraints, the options considered, and why one path won.

“A demo proves possibility. Production architecture proves responsibility.”

Writing on these themes is published as the work allows. Until then, the fastest way to hear the reasoning is a conversation.

Bring a decision to work through

Your startup architecture partner

I help early teams find the path before the playbook exists.

My strongest body of work is with founders and lean engineering teams making product, model, cloud, evaluation, and cost decisions while the product itself is still moving.

Through Microsoft for Startups, I see the patterns between a promising demo and a production system across many teams. That pattern recognition is the leverage I bring to your table. You work directly with me, with no sales layer or junior handoff.

2022

Production AI before the boom

Joined OneReach.ai in October 2022 and began building conversational AI and orchestration systems at production scale.

2023-24

One-person platform ownership

Owned roadmap, tooling, and architecture for Wind River's Engineering Excellence function before moving into full-stack platform engineering.

Today

Startup architecture at the decision table

Work with founders and lean engineering teams on AI architecture, agent systems, evaluation, and technical direction through Microsoft for Startups.

Also

Founder and operator

Founded Presencia Studio and operate production systems personally, which keeps the advice honest about what happens after launch.

Current independent work is limited to non-conflicting engagements and does not imply endorsement by Microsoft, Accenture, or any current or former employer.

Start a conversation

Two practices, one point of view

Advisory is mine. Engineering capacity is Presencia's.

I advise startup teams personally: the decision, the architecture, the sequence. Presencia Studio is the engineering organization I founded to build and operate software and AI systems for businesses.

Running production systems keeps the architecture work grounded in the decisions teams actually live with after launch. When an engagement needs more implementation capacity than advisory provides, Presencia can become relevant — but it is never the default.

Judgment

Yenson Umaña

What should we build, how should we build it, and what needs to be true before we commit?

Engineering capability

Presencia Studio

What is holding the business back, and what system should we build to solve it?

Before we talk

Useful answers, without the sales call.

Do you implement the systems you design?

I use prototypes and technical validation when they reduce a material risk, and I can support an internal team through delivery. The core offer is senior judgment — direction, architecture, and adoption — not open-ended outsourced development.

What kind of company is the best fit?

Startups and scale-ups with real momentum, a founder or CTO close to the decision, a team ready to execute, and an AI initiative consequential enough to shape the next year. If what you need is implementation capacity rather than technical judgment, I will say so early.

How is this different from what Presencia Studio does?

I advise founders and engineering teams personally on decisions and architecture. Presencia Studio is the engineering organization I founded to build and operate systems. They reinforce each other, but an advisory engagement never assumes Presencia does the building.

How does this work alongside your current role?

I accept a limited number of non-conflicting engagements, each subject to a conflict review. Client work is independent and does not imply endorsement by Microsoft, Accenture, or any current or former employer.

Can you work with a distributed or international team?

Yes. I work from Costa Rica with teams globally in English, using a mix of live working sessions and documented asynchronous decisions.

How do you handle confidentiality?

A mutual NDA is available on request. Engagement boundaries, data access, model usage, IP ownership, and retention expectations are agreed before sensitive material is shared.

Free 30-minute discovery

Bring the AI problem your startup cannot yet turn into a path.

Bring the unresolved decision, the competing technical paths, or the prototype that has to become production. It is a technical conversation: we clarify the real constraint and the next decision, whether or not we end up working together.

Book a discovery call