AI engineering · AI-native applications
Most AI roll-outs change the tool, not the work.
The common shape: a tool is bought, seats are assigned, usage is reported upward as success, and the process around the work is identical to the one that existed before. The return on that is close to zero, and everybody involved can tell. The gains come from redesigning the process. The tool is the cheap part.
The premise
A faster version of a step that should not exist is not an improvement.
This is not scepticism about the technology. We build with these tools every day and they are the reason we can rebuild a licensed product in weeks. It is scepticism about a particular way of buying them, which is to treat adoption as a distribution problem — get the licence to everyone — rather than as a process problem.
A process that exists in an organisation is usually shaped by a constraint. Three approvals because drafting was expensive. A weekly batch because the report took two days to compile. A handover to a specialist team because only they had the context. Put an assistant against that process unchanged and you get the same shape, slightly faster, with a new subscription attached.
Remove the constraint properly and the shape changes: the approvals collapse, the batch becomes continuous, the handover disappears. That is where the return is, and it is organisational work rather than technical work, which is why it is the part that gets skipped.
So we sequence it deliberately. Assessment first, because you cannot redesign a process you have not measured. Tools second, chosen against your real work. Redesign third, on two or three workflows rather than twenty. Enablement fourth, for a full quarter, because habits are slower than software. Measurement throughout, on numbers that can go down as well as up.
The distribution mistake
A target of 80 percent of staff licensed by the end of the quarter. Achievable without changing anything, which is exactly why it is the wrong target.
The pilot mistake
A six-week pilot with volunteers who were already enthusiastic, producing a number that will not reproduce anywhere else in the organisation.
The platform mistake
Eighteen months building an internal platform, by which point the underlying models have changed twice and the platform is the constraint.
The measurement mistake
Reporting prompts issued and messages sent, because they are easy to count and they only ever go up.
The programme
Five stages, and you can stop after any of them.
Each stage produces something usable on its own. Plenty of clients buy the assessment and the tool roll-out, run the process work with their own people, and come back for enablement a quarter later. That is a legitimate shape and the pricing is arranged so it is possible.
Assessment
3 to 4 weeksBefore choosing a tool, find out where the work actually is. We sample real work rather than asking people to estimate it, because self-reported time allocation is consistently wrong in the same direction. The output is a ranked list of workflows and an honest read on which of them the current generation of models can move.
Deliverables
- A work inventory across the functions in scope, built from work sampling rather than surveys
- A ranked shortlist of 8 to 12 candidate workflows, each with a measured baseline cycle time
- A data and access readiness assessment: what exists, where it sits, what classification it carries
- A capability read per workflow — what today’s models do well, do badly, and cannot do at all
- A written list of workflows we recommend you do not attempt this year, with reasons
Tool selection and roll-out
4 to 6 weeksEvaluation against your real work, not against vendor demos. Every assistant looks excellent on a curated example. The point of this stage is to find out how the shortlist performs on the three workflows you actually ranked highest, with your data shape and your access constraints in the way.
Deliverables
- A structured evaluation against your own workflows, scored, with the test set kept so it can be re-run
- Security and data protection review, including where prompts and outputs are retained and by whom
- Procurement position: contract terms, data processing agreement, and the clauses worth arguing about
- Identity integration, so access follows your joiner-mover-leaver process rather than a separate admin console
- A licence right-sizing model, because the standard failure is buying an enterprise seat for everyone in month one
- Deployment to the first cohort, with a support path that is a real one
Process redesign
6 to 10 weeks, two or three workflows in parallelThe stage that produces the return, and the one most programmes skip. Adding an assistant to an unchanged process gets you a faster version of a step that should not exist. We take the workflow apart, remove the steps that only existed because the work used to be slow, and rebuild around what the tools are genuinely good at.
Deliverables
- A current-state map per workflow with measured cycle time, handoffs and rework points
- A redesigned workflow — steps removed, steps merged, steps automated, with the reasoning written down
- Working automation for the parts worth automating, in your systems, owned by you
- Updated controls and approvals, agreed with risk and compliance rather than discovered by them
- A before-and-after measurement, published internally including where it did not improve
Enablement
8 to 12 weeks, overlapping stage 3Adoption is a habit problem. Habits take a quarter, not a workshop. The organisations where this sticks are the ones with a credible person inside each function who uses the tools daily and is available when a colleague gets stuck at four in the afternoon.
Deliverables
- A named champion per function, trained deeply, with time formally allocated to the role
- Role-specific training built on your redesigned workflows, not on generic prompt technique
- An internal pattern and prompt library in your own systems, with owners and a review cadence
- Open office hours for the first 8 weeks, then a handover of that slot to your champions
- A written escalation path for the failure modes: wrong output, sensitive data, tool outage
Measurement
Ongoing, monthly reportingMost AI programmes are reported on inputs because inputs are easy to count. We report on whether the work got faster and whether it had to be redone, per workflow, against the baseline taken in stage 1. Where a workflow did not improve, that is reported too, with what we think went wrong.
Deliverables
- A monthly board-level report: cycle time, rework rate, and workflows actually changed
- Per-workflow measurement against the stage 1 baseline, with the sample size stated
- A quarterly review of the workflows we said no to, because model capability moves
- A cost line: licence spend, our fee, and the measured return, with the assumptions shown
The proof
We publish a reference application every day. That is the evidence.
We would rather not argue about AI productivity in the abstract, so here is the thing we actually do with it. Conseiltek publishes one open reference application every day — 37 in the catalogue, 3 published so far — each with runnable code, an AWS architecture with Terraform, an Azure architecture with Bicep, a parity matrix against the product it replaces and a cost model. That output is only possible because the delivery process was rebuilt around these tools rather than having them added to it.
Specification before code
Every application starts as a parity matrix and a data model, written and argued over with model assistance, before anything is generated. The specification is the artefact that makes the speed reproducible.
Reference implementations, not blank pages
Each new application inherits an architecture, an infrastructure baseline and a deployment path that already work. The novel part of any given build is a minority of it, and that is by design.
Tests and migration tooling first
The parts of engineering that are tedious and well-specified are the parts these tools are best at. We spend the saved time on the parts that need judgement.
Human review on everything shipped
A senior engineer owns every published repository by name. Generated code that nobody understood well enough to defend does not go in.
The process changed, not just the toolbar
Our review cadence, our branching model and our definition of done are all different from what they were two years ago. That is the actual adoption story.
Published so you can check
Apache-2.0 on GitHub, with the cost models and the honest parity rows included. If the claim were inflated, the repositories would show it.
Browse Techtons → · the same engineers who build it are the ones who run the adoption programme
Measurement
Four numbers, and five we refuse to report.
Every measure below is taken against a baseline captured in stage 1, before anything changed. Sample sizes are stated. Where a workflow got worse, it is in the report, because a measurement framework that can only produce good news is not a measurement framework.
Cycle time, per workflow
Start to finish, measured against a baseline taken before anything changed. Per workflow, because an organisation-wide average hides the two that worked behind the six that did not.
Rework rate
How often output goes back for correction. The failure mode of a fast, poorly designed AI workflow is more output that needs more checking, and rework is what catches it.
Workflows changed
A count of processes that are genuinely different from how they ran before. This is the number the programme is actually accountable for, and it is usually small and honest.
Time to competence
How long a new person in a role takes to reach normal productivity. It moves late, and it is the strongest signal that the change has become part of how the work is done.
What we will not report as adoption
- Seat count, or percentage of staff with a licence
- Messages sent, prompts issued, or sessions per user per week
- Lines of code accepted from an assistant
- Self-reported time saved, collected by survey
- Attendance at training
All of these correlate with spending money and none of them correlate with the work being different. They are worth tracking operationally, for licence right-sizing. They are not the programme's result and we will not present them as one.
Governance and data
Decide this once, in writing, before the first licence.
Governance that arrives after roll-out is remediation. Six decisions cover most of it, and none of them are hard — they are simply unowned in most organisations until something goes wrong.
Data classification, before tool selection
Which classifications may enter which tool, decided once and written down, rather than negotiated per team per week. Most organisations already have the classification scheme; almost none have mapped it onto AI tooling.
Retention and training terms in writing
What the vendor retains, for how long, whether it is used for training, and which jurisdiction it sits in. These terms differ between a vendor’s consumer tier and its enterprise tier, and the difference is usually the whole argument.
Access through your identity provider
Provisioning and deprovisioning through your joiner-mover-leaver process. An AI tool with its own user list is an orphaned account waiting to happen, and it will not appear in your next access certification.
A human accountable for every output that leaves
Not a review checkbox — a named role that owns the result. Where the output goes to a customer, a regulator or a court, the accountability model matters more than the model.
Logging you can audit
Prompt and output logging where classification requires it, retained under your own retention policy, queryable by your security team without a vendor support ticket.
A written position on shadow usage
People are already using these tools on their own accounts. A policy that pretends otherwise produces an unlogged version of the same risk. The workable answer is a sanctioned path good enough that the unsanctioned one is not worth the effort.
Where the systems in scope are ones we also build or operate, this work joins up with our security and compliance and identity lifecycle practices rather than running as a parallel exercise with its own committee.
Where AI is not the answer yet
Four kinds of work we will tell you to leave alone.
Yet is doing real work in that heading — model capability moves, and we review this list with clients every quarter. But a programme that attempts everything produces four visible failures and loses the mandate for the things that would have worked.
Work where being wrong is expensive and checking is as slow as doing
If verifying the output takes as long as producing it, there is no gain, only a different kind of labour. Legal drafting against unusual precedent, safety-critical specification and regulatory filings frequently sit here.
Anything that depends on knowledge nobody wrote down
If the context lives in individual inboxes, eleven versions of a spreadsheet and one long-serving colleague’s memory, the first project is a knowledge project. We will tell you that rather than disguise a data programme as an AI programme.
Long-horizon autonomous work with no checkpoints
Multi-day, multi-step tasks with no human in the loop remain unreliable in ways that are hard to detect until the end. We design for checkpoints, and where checkpoints are impossible we do not automate it yet.
Judgement calls the organisation has not made itself
Where the real problem is that two directors disagree about the policy, a model will produce a confident answer to a question your organisation has not resolved. That is worse than slow.
The assessment produces this list for your organisation specifically, and it is a deliverable rather than a caveat. Being explicit about the workflows we are not attempting is what makes the ones we do attempt credible — and it is the same reason every application in Techtons carries parity rows where the honest answer is no.
Start with the assessment. Three to four weeks, fixed price.
We sample how the work actually runs across the functions in scope, rank the workflows by what current tools can genuinely move, measure a baseline you can be held to later, and write down which workflows we recommend you leave alone this year. You keep the report whether or not the programme continues.