Skip to content

AI talent

Prompt and LLM Engineer

Designs, tests and maintains the prompt and context layer, and holds output quality over time.

The role

What the job actually is.

This is the role that owns what the model is given and what comes back. It covers prompt structure, how retrieved context is assembled into a request, and the test sets that say whether the result is working. The word most people miss is maintains: a prompt is not finished when it works, because the model underneath it will change and the documents it draws on will drift. A good LLM engineer treats prompts as versioned artefacts with regressions attached, and can tell you whether a bad answer came from the prompt or from the retrieval layer before touching either. Where the retrieval layer itself is designed and built, that is the AI Data Engineer's work, and the two roles are normally hired together.

What we screen for

The questions that separate the field.

Asked by someone who has built the thing, and designed to catch this role's specific failure rather than to confirm a general impression.

  1. A worked example of an output that got worse after a model change, and what they did about it.
  2. Whether they version prompts, and what the diff of a prompt change looks like to them.
  3. Whether they can distinguish a retrieval problem from a prompt problem, and how they proved which it was.
  4. How they assemble retrieved context into a request, and what they cut when it stopped fitting.

The common mis-hire

The one you have probably already made.

Someone who is fluent in a chat interface. Fluency is not engineering, and it interviews extremely well because the demonstration is the skill. Ask for the regression they caught rather than the prompt they are proud of.

In the estate

Where this role works, and what we screen it against.

The layers this family works at, lit. These are the tools we screen against. Naming one says we can test for it, not that we have delivered on it.

Evaluation and observability

Spans every layer. Without it a system is shipped on impressions.

  • LangSmith

Also screened against

  • Langfuse
  • Ragas
  • DeepEval
Role families we place here

Experience and delivery

Copilots and agents inside a business process, and the interaction design that makes an uncertain system usable.

Orchestration and agents

Where an agent's steps, tools and state are defined, and where its failures are caught before a user meets them.

Also screened against

  • LangChain
  • LangGraph
  • Pydantic AI

Models

The models themselves, and the platforms an enterprise hosts them through.

  • Anthropic (Claude)
  • OpenAI
  • Google (Gemini)
  • Mistral
  • Cohere
  • Azure AI Foundry
  • Google Vertex AI
  • Hugging Face

Also screened against

  • Meta (Llama)

Data and grounding

What the model is grounded in, and the integration work that gets enterprise data to where it can reach it.

  • Azure AI Search
  • Pinecone
  • Elasticsearch

Also screened against

  • pgvector
  • Weaviate
  • Qdrant
  • LlamaIndex

Systems you already run

The seven platform desks Yallo staffs. Almost no AI work is greenfield; it lands here.

Role families we place here

Governance, risk and safety

Spans every layer. Named as what governance roles are screened against; what any of them obliges is your counsel's call.

  • EU AI Act
  • ISO/IEC 42001
  • ISO/IEC 23894
  • NIST AI Risk Management Framework
  • OWASP Top 10 for LLM Applications
Role families we place here
Naming a technology here says we screen against it, not that we have delivered on it. The role families on each layer are the ones we place there.

Seniority

What changes between mid, senior and lead.

The grade is a description of what the person owns, not a band. Rates come with the shortlist.

Mid
Writes and tunes prompts against an existing evaluation set and retrieval design.
Senior
Owns the prompt and context layer end to end, and builds the evaluation that decides whether a change ships.
Lead
Sets the prompt and context standard across teams, owns model-change response, and holds quality across releases rather than at one point in time.

In a programme

When this role is needed, and what blocks it.

Starts once there is a grounded corpus to write against, which puts it a step behind the AI Data Engineer rather than alongside it. The dependency that bites is the same one: this role is blocked by whatever governs the documents, not by procurement. Keep it retained after go-live, because the maintenance limb of the job only begins there.

Ask

Send the brief, get a screened LLM Engineer shortlist.

Tell us the programme, the stack and the timeline.