Skip to content

Fractional Teammates · AI & Automation

A generative AI developer for the part after the demo

Retrieval that grounds answers in your own content, an evaluation set that catches regressions before release, guardrails, and a named person who signs off. A senior OlDevs developer does this inside your repositories, on a set share of the week.

A senior specialist, not a junior placement You own the work and the accounts Reply within one business day

AI

Copilot · fine-tuned on your data

Live
Which invoices are overdue by more than 30 days?
14 invoices totalling 42,300 are overdue. Three accounts carry 68% of it — I’ve drafted reminders for review.

0.94

F1 score

86ms

Inference

12k/d

Requests

What a generative ai developer is

OlDevs fractional generative AI developers build and maintain the software around a language model.Retrieval that grounds answers in your own content, prompt and evaluation harnesses that score a change before it ships, guardrails for unsafe or off-topic output, and control over token cost and latency. They also write down who signs off on customer-facing output, and how a wrong answer gets reported and fixed. The seat is a set share of a senior week, worked inside your repositories, tickets and standups. Take it when the LLM work keeps arriving but the demo is still the only thing that has shipped.

Key facts

01Engagement
Ongoing fractional, quote-based
02Commitment
A set share of a senior week
03Works in
Your repositories, tickets and provider accounts
04Ownership
You own the code, prompts and evaluation sets
05Typical commitment
1–2 days a week
06Judged on
Accuracy on your own eval set · Cost per request · p95 latency

AI & Automation

What this developer builds and maintains

What this teammate takes off your plate

01

Retrieval that grounds answers in your own content

Builds the ingestion, chunking, embedding and search layer so the model answers from your documents rather than guesswork. Handles permissions, freshness and the awkward source formats nobody wants to own.

RAG · Embeddings · Chunking · Access rules · Re-indexing

02

Prompt and evaluation harnesses

Turns prompts into versioned assets with a test set behind them. Every change is scored against saved cases, so you can see whether a release improved behaviour or quietly broke it.

Golden sets · Regression runs · Scoring · Versioning · LLM-as-judge

03

Guardrails and safe failure

Defines what the feature must refuse, how it escalates to a person, and what it says when a model is unavailable. Adds input filtering, output checks and logging you can audit later.

Refusals · Input filtering · Fallbacks · Audit logs · Escalation

04

Cost and latency control

Measures tokens, cache hits and response times per feature, then tunes model choice, context size and caching against them. You get a monthly picture of which features consume the most and why.

Token budgets · Caching · Model routing · Latency · Batching

05

The review process behind customer-facing output

Writes down who approves what, where the human sits in the loop, and how errors are reported and fixed. Includes accessibility checks to WCAG 2.2 AA on anything users read.

Human in the loop · Sign-off · Incident path · WCAG 2.2 AA · Disclosure

06

Shipping it the way you ship everything else

Deploys the feature through your existing pipeline, with monitoring, secrets handling and a rollback. Documents it so your own developers can change a prompt or swap a model without calling us.

CI/CD · Monitoring · Secrets · Rollback · Runbooks

Benefits

What changes when an LLM feature has to ship

01

Hard model calls stop being guesses

Which model, whether to retrieve or fine-tune, how much context is enough, when a prompt is finished: these come in bursts and get decided by someone who has made them repeatedly, on a workload that would not fill a permanent role.

02

Quality stops being a matter of opinion

Instead of three people reading the same answer and disagreeing, changes get judged against a set of cases everyone has agreed on. Product, support and legal can point at the same evidence when they push back.

03

Provider changes stop being emergencies

Model providers deprecate versions and change behaviour on their own schedule, not yours. With a baseline to compare against and one place where the model is chosen, that becomes scheduled work rather than a scramble.

04

The feature reaches customers, not just staff

Plenty of prototypes stop at internal users because nobody will sign off on what a model might say to a paying customer. When the failure modes are named and bounded, that sign-off conversation has something concrete to weigh.

05

Bad AI ideas get stopped early

Not every request suits a language model; some are a search problem, a form, or a rule. Saying so early, with reasons your team can check, keeps months of work from going into a feature that cannot be made reliable.

06

Your own developers can take it over

Prompts, evaluation sets and the reasoning behind each retrieval choice live in your repository, written for whoever reads them next. Nothing load-bearing about the system sits only in one person's head or leaves when the engagement does.

How it works

How we get to something you can ship

  1. What happens if it is wrong

    The first conversation is about the use case, the data behind it and the cost of a bad answer. If a fractional seat is the wrong shape, we say so. Then we quote.

  2. Your accounts, not our sandbox

    The developer works in your repository, ticket tracker, chat and model provider accounts, on your own licences. Nothing important is built somewhere you cannot reach after we leave.

  3. A baseline before any prompt changes

    Real cases with expected answers, collected from your own history and your support queue. Without that baseline, every later claim that behaviour improved is somebody's impression.

  4. Narrow first, behind a flag

    The first feature goes to a small internal group before any customer sees it. We would rather show something narrow that holds up than a broad demo that does not.

  5. A release gate with a name on it

    Nothing customer-facing ships without an evaluation run and a named approver from your side. The scores, the failures and what is queued next go in your tracker each week.

Who it's for

Teams this seat is built for

Continuing LLM work needs someone who owns evaluation and guardrails between releases. That is rarely a full week of work, and rarely something a busy in-house team gets to.

A prototype that convinced the room

The demo won the argument. What it now needs is retrieval that stays current, an evaluation gate before each release, and a path for the answer that goes wrong overnight.

Software adding an LLM feature to a paid product

Your customers already rely on the product, so an assistant that invents an answer is a support problem and a refund conversation, not an interesting failure.

In-house developers with no LLM specialist

Your team is capable and fully booked. The teammate carries retrieval, evaluation and guardrail work, and writes it so your developers can maintain it once the seat ends.

Answers that are published, not just used

Associations, government bodies and regulated firms, where a wrong public answer becomes a complaint. Review, records, bilingual output and accessibility have to be in the first version.

12+

Years of studio experience since 2014

1–2

Typical commitment — 1–2 days a week

EN/FR

Languages this engagement can be delivered in

FAQ

Questions we get before an LLM project starts

No, and be careful with anyone who does. Language models generate plausible text; they do not know facts. What we can do is reduce the rate and contain the damage: ground answers in your own content, cite sources in the interface, measure accuracy against a saved test set, refuse rather than guess outside scope, and route anything sensitive to a person before it reaches a customer.

They were never ours to keep. Code and prompts live in your repository, evaluation sets sit beside them, and the provider accounts are in your name on your own licences. At the end you get a walkthrough, written runbooks and a final pass that removes our access. The evaluation set is the part people forget, and it is what lets the next developer change anything safely.

Choosing a provider is your decision; we bring the comparison, the data-residency implications and the migration effort in writing. In practice that means one of the major hosted models, with an open-weight model considered where data cannot leave your environment. We build retrieval, evaluation and guardrails as your own layer, so the model underneath can be swapped without rewriting the feature.

When the LLM feature is the product itself — that is a team, not a share of one week. When you want a fixed scope with a firm end date, which is a project and should be bought as one. When nobody on your side will own what the model is allowed to say, because that decision cannot be outsourced. And when the honest answer is that plain search would serve your users better than a model.

The first weeks go into a baseline: a real evaluation set, a look at your data, and a written view of where the current behaviour fails. Something running usually follows soon after, behind a flag and in front of a small internal group first. We will not put a date on a release before the baseline exists, because until then nobody can say what better means.

Usually the order of work. A developer without LLM experience tends to ship prompt changes and hope; the teammate puts a measurement in front of each change, and takes on the parts that get skipped — retrieval freshness, refusals, logging, cost. They review your developer's pull requests rather than working around them, and the aim is that your team keeps the feature after the seat ends.

Still have a question? Ask us when you request a quote

Let’s connect

Let’s talk about the generative ai developer gap.

Tell us what is not getting done and roughly how much of a week it needs. We’ll reply within one business day with who would cover it and a tailored quote — no obligation.

We’ll only use your details to prepare your quote. No lists, no spam.

Call us Request a quote