Agentic AI

What a forward deployed engineer does, and how to tell if you are ready

  • Judge a forward deployed engineer, or FDE, by what still runs after handover, measured on the client's own numbers.
  • Production code inside the client's systems is the line between an FDE and the roles next to it.
  • The work before the build matters most: real production data, a one-page scope and a baseline to beat.
  • Take the 12-question readiness test below, then start on your two lowest scores.

We judge a forward deployed engineer, or FDE, by what still runs in the client's operations after handover, measured on the client's own numbers. Demos and deals won are the wrong measure, because the job exists to close the gap between pilot and production. Gartner predicts that over 40 percent of agentic AI projects, in which AI takes actions on its own, will be canceled by the end of 2027.

The job in one paragraph

An FDE is a software engineer who works inside a client's organization until a product runs in its real operations. At extendfuture, the FDE takes AI systems the last mile, from a prototype on sample data to a system that runs the client's daily work.

The FDE fits AI models to each client's data and software, then owns deployment, rollout, adoption and a clean handover. The FDE also translates both ways: business problems into technical scope, and technical reality back to the people who decide.

Where the role started

Palantir's original definition still holds, and we use it. Palantir created the role in the early 2010s, and Tom Hollands of Andreessen Horowitz, or a16z, dates it to 2011.

A 2019 Palantir post calls its FDEs Deltas. It contrasts a product engineer's "one capability, many customers" with a Delta's "one customer, many capabilities." When a missing feature blocks a client's goal, Deltas sometimes write the core product code.

Who hires FDEs now

On 10 October 2026, OpenAI's job board listed roles with "forward deployed" in the title in 14 cities, and Databricks listed more than a hundred. OpenAI's openings include industry teams for healthcare, legal work and financial services. Anthropic lists FDE roles in London, Munich and Paris. Palantir still hires under the title, and Google Cloud was hiring FDEs for a new AI organization in May 2026.

a16z partner Joe Schmidt argued in 2025 that AI startups should accept lower margins on hands-on deployment, because it helps win larger contracts. We agree only when each deployment leaves parts the next one can reuse. Without them, every deployment costs as much as the last.

How the FDE differs from nearby roles

Production code inside the client's systems is the line. We would not give the title to a role that does not write it.

Anthropic's listings show the split inside one company. Its Technical Deployment Lead owns scope and value measurement, and "won't write production code." Its FDE builds production applications inside client systems.

The FDE and the roles next to it

Sales engineerSolutions engineer or architectConsultantProduct engineerFDE
Works onWinning and renewing salesAdoption, starting before the saleA defined questionOne part of the vendor's productA single client's operations
DeliversProduct presentations and technical sales supportAdvice, prototypes and examples the client extendsA one-time analysis or recommendationOne capability for many clientsProduction code in the client's systems, handed over
SourceUS Bureau of Labor StatisticsAnthropic listingPalantirPalantirOur definition, above
  • Sales engineer

    Works on
    Winning and renewing sales
    Delivers
    Product presentations and technical sales support
  • Solutions engineer or architect

    Works on
    Adoption, starting before the sale
    Delivers
    Advice, prototypes and examples the client extends
  • Consultant

    Works on
    A defined question
    Delivers
    A one-time analysis or recommendation
    Source
    Palantir
  • Product engineer

    Works on
    One part of the vendor's product
    Delivers
    One capability for many clients
    Source
    Palantir
  • FDE

    Works on
    A single client's operations
    Delivers
    Production code in the client's systems, handed over
    Source
    Our definition, above

Responsibilities before, during and after a deployment

We weight the work before the build most, because several of the failures below start there.

Before, during and after a deployment

  1. beforeMap the workflow and the system that holds the official records

  2. beforeGet real production data early

    Curated samples leave out the cases that break a model.

  3. beforeWrite a one-page scope

    What is in, what is out, the success measure and the baseline number to beat.

  4. beforeAgree what data the system may touch

    And how it passes the client's security and compliance review. This post is not legal advice.

  5. duringIntegrate with the client's sign-in system, data sources and release process

  6. duringBuild an eval set and rerun it after every change

    An eval set is a fixed list of real cases with known correct answers.

  7. duringPut a human check on decisions where a wrong answer is expensive

  8. duringRoll out in stages

    Train users, track adoption and trade scope against the date in writing.

  9. afterHand over a runbook

    The written guide to running and fixing the system, with alerts and a named client owner.

  10. afterReport the result against the baseline

  11. afterTurn custom work into reusable parts

    As OpenAI's and Anthropic's listings ask.

Skills

We would weigh the last three skills as heavily as the engineering, because they decide whether the work survives handover.

  • Full-stack engineering in Python or TypeScript, meaning interface, server and database, plus speed in code you did not write.
  • Breaking a business problem down to the code that solves it, which a Palantir Delta calls "technical decomp."
  • Production experience with language models, eval sets and monitoring.
  • A record of shipping inside someone else's environment, under their access rules.
  • The maturity to hold a room with client engineers and their executives.
  • Judgment about when a fix belongs in the product and when it stays custom.
  • Clear writing for clients in other time zones.

A typical week

Expect less coding than the title suggests. Gergely Orosz of The Pragmatic Engineer puts the split at roughly a quarter coding, half integration work and a quarter meetings and client support. OpenAI's San Francisco listing expects up to 50 percent travel.

On an AI deployment, we would start the week with the weekend's traces, the step-by-step logs of what the system did, and end it with a written update to the sponsor.

How success is measured

Palantir says Deltas measure success by impact on the client's goal. OpenAI's FDE listing names "production adoption, measurable workflow impact, and eval-driven feedback." We count success only after handover, when four things are true.

  1. The system runs in production on real data, used by the people it was built for.
  2. A workflow number, such as accuracy or cost per outcome, moved against the baseline.
  3. Someone other than the FDE runs it from the runbook.
  4. Part of the work is reusable on the next deployment.

Common ways FDEs fail

In our view, each common failure has a habit that prevents it.

  • The FDE scopes on samples, so the demo breaks on production data. Build the eval set from real data first.
  • The FDE says yes to every date. Trade scope, never quality, in writing.
  • The FDE builds one-offs. a16z's Marc Andrusko relays the warning that this leads to "thousands of bespoke deployments that are impossible to maintain or upgrade." Decide for each fix whether it belongs in the product.
  • The FDE advises instead of shipping, the line Palantir draws between Deltas and consultants.
  • The FDE launches without a baseline, so nobody can show what changed.
  • The FDE leaves no runbook or owner, so nobody can fix the system when it fails.

The readiness test

Are you ready to work as an FDE?

Score your record, not your confidence: 2 if you have done it and could show evidence, 1 if partly, 0 if not.

1Have you shipped a change, in your first week, to a codebase you did not write?

2Have you deployed into an environment you did not control, such as a client's cloud account?

3Have you changed a plan before launch because production data differed from the scope?

4Have you built and hosted a small tool alone, with an interface, a server and a database?

5Have you shipped a language model feature and measured it with an eval set?

6Has anyone signed off a one-page scope you wrote, with a success measure?

7Have you told a senior stakeholder "not by that date" and kept their trust?

8Have you explained a technical trade-off to an executive and left with a decision?

9Has someone else run your system using only your handover document?

10Did you record a baseline before your last launch and report the change after?

11Have you shipped after weeks with no one setting your next task?

12Has a fix you built for a single client or team been reused by the next?

12 questions. Answer yes only for what is true today.

A score of 20 to 24 means you can do the job now, and 13 to 19 means you are close. At 12 or less, build the base in your current role first.

One scenario question

Answer before reading the checklist.

You are three weeks into deploying an AI assistant that drafts replies for a client's support team. On Wednesday, the compliance lead finds drafts quoting a refund policy the client retired last month, and asks you to stop. The executive sponsor wants a second team live on Monday. What do you do, technically and with the people, before Monday? What would you cut, what would you keep, and how would you say no?

What a strong answer covers

  1. Use the logs to find every affected draft

    Give the compliance lead the facts.

  2. Fix the cause, rather than patching the prompt

    Such as a stale document in the knowledge base.

  3. Add retired-policy cases to the eval set

    Rerun it before anything ships.

  4. Go live on Monday only where it is safe

    Such as ticket types that involve no policy.

  5. Keep the human check on refund replies

    And the audit log and the rollback plan.

  6. Meet the compliance lead and the sponsor together

    Offer options with risks, and record the decision.

  7. Pair the no with a path

    Such as "Order-status tickets on Monday, refunds once the new eval cases pass."

A weak answer promises everything by working the weekend, or refuses with no alternative.

How to close the gaps

Start with your two lowest scores.

One move for each gap

  1. Get a fix merged into an unfamiliar open-source project

    For questions 1 and 4.

  2. Deploy an internal tool for another team inside their accounts

    For question 2.

  3. Build a small assistant on a public dataset and track its eval score per change

    For questions 3 and 5. Use the open-source tools promptfoo for evals and Langfuse for traces.

  4. Get a one-page scope for your next task signed by the decision maker

    For questions 6 to 8. When a date slips, offer a smaller scope by that date.

  5. Have a colleague run your system from your runbook

    For question 9. Fix every gap.

  6. Record a baseline before your next change

    For questions 10 to 12. Take on a problem nobody owns, and ask after each fix whether the next team needs it.

To build that record on production AI systems, see extendfuture's open roles.

Sources

· views
Amol PatilFounderFounded extendfuture in 2019. Has shipped computer vision, voice and agentic systems into production across ten industries. Amol Patil on LinkedIn

Working on something in this territory?

Tell us what you are trying to win. We answer within one business day, from the people who build.