BACKGROUND
Estimating a custom software build is hard. The number you put in a proposal usually gets locked in, and then the delivery team has to build to it. So everyone from the buyer to the people doing the work needs to trust it: it has to win the bid, and it has to be a target the team can actually hit.
AI adds a new pressure on top of that. Buyers now expect it to take real money off a quote, and some arrive expecting a project to cost a fraction of what it used to. Nobody, on either side, can point to which of those savings are real. AI might make one task much faster without changing the shape of the whole project, and an invented reduction that sounds confident can do real damage. It’s the delivery team that pays for it later, building against hours that were never real.
We’ve been running an internal effort at OXD to work AI into our own delivery, partly to go faster, partly to get better at working with it. That work starts with people researching each task and validating by hand how far its hours can reasonably come down. That manual validation is the foundation the AI estimation is built on.
The Baseline app is the next step in that work. It sits between the internal AI experiments the team runs and the estimates that go into real projects, turning what those experiments find into numbers grounded in past delivery. It’s built for estimating custom software builds, where the estimate becomes the budget the team has to deliver against.
HYPOTHESIS
As AI enters delivery, nobody really knows how much it changes a budget. The answer is uneven. It speeds up some tasks a lot and others barely, and the pattern shifts as the tooling improves.
That effect can be measured task by task, bounded by what hands-on research actually found, and turned into an estimate that shows where AI moves the work and where it doesn’t. As real projects get logged, the picture sharpens, so vendors and buyers alike get a grounded read on how AI is affecting estimation.
APPROACH
Baseline turns a project into a costed estimate, with a confidence rating on every saving. An estimator works through it like this:
- Pick a project type. Baseline pulls in a standard task list, each task priced from what similar past projects actually took.
- Shape the list to the job: add, remove, or reorganize tasks.
- Ask AI where the hours can come down. It proposes a reduction on each task, held inside a range that research has already set for it.
- Review each suggestion, with its rationale and the tradeoff it names, and accept or reject it.
- Export a costed estimate: hours, cost, and a confidence rating on every saving.
Starting from real history
Every task starts at what past projects actually took, with no AI involved. The hours come from logged projects of a similar size, type, and sector, taking the middle of what those projects ran to as the starting number, with a realistic low and high around it. That gives the estimate a factual floor. The model never sets the base number itself.
When too few past projects match, the tool flags the number as an estimate and uses a rough default. Early on that happens a lot. As real projects get logged, more tasks move onto real data and the starting numbers firm up.

Setting the limits before the AI runs
Baseline lets AI propose where a task’s hours can come down. To keep those proposals grounded, a person sets the limits before the model runs. Someone studies the task, drawing on a hands-on experiment, published research, or our own delivery data, and records how far its hours can realistically come down as a range with the evidence attached. This is the manual research from the internal experiment, and it’s what the AI works within. Sometimes the honest answer is that the task can’t come down at all, and that gets recorded just as deliberately.

That range is general on purpose. It’s the same for every project that uses the task, and on its own it can’t say where a given task should land this time. A task that could come down 30 to 50 percent might warrant the full amount on a fresh build and nothing on a legacy integration.
That placement is what the Ask AI step is for. For a task, it reads the actual project, the brief, the other tasks, and any note the estimator adds, then proposes where inside the researched range the task should sit, or whether it should move at all, with a rationale and a tradeoff. If it reaches past the ceiling, the tool pulls the number back and marks that it did. The research says how far a task could ever come down; the model works out how far it should on this project, and a person makes the final call.

Showing the confidence
A saving is only as useful as the evidence under it. Every one is labeled evidenced, plausible, or speculative, and the estimate shows how much of the total sits in each, with a warning when too much leans on the weak end. So anyone reading it, vendor or client, can see how much of the number rests on solid ground and how much is still unproven. Where the evidence is thin, that saving is easy to set aside and fall back to the base figure.

TAKEAWAYS
01
Every estimate starts from what past projects actually cost. The AI can only adjust that number; the base always comes from real delivery. So the price a buyer sees has history behind it, and it stays a target the team can actually hit. It’s the difference between a number you can defend and one you can’t.
02
AI can only cut a task as far as the research says it can. Before the model runs, someone works out how much AI realistically saves on that task and records the range with the evidence behind it. The model works inside that range and gets pulled back if it reaches past it. So every saving reflects what a person actually found AI could do on that task.
03
Every saving comes with a rating for how solid the evidence is. Savings are labeled evidenced, plausible, or speculative, and the estimate shows how much of the total rests on each. Anyone reading it can weigh the confident savings and the shaky ones for what they are. It’s also what turns a vague expectation, like AI knocking 70 percent off a project, into a grounded picture of where the savings really are.
04
The AI is kept deliberately narrow. It only weighs in where a person would actually need judgment: reading a brief, proposing a reduction, writing the summary. Everything else is ordinary calculation on known numbers. Keeping it small means every reduction traces back to a specific piece of evidence, so no number goes in without a reason behind it.
05
The approach outlasts any one tool, including AI. AI is just the first efficiency run through it. Because completed projects feed their real hours back in, and the researched limits get revised as that reality lands, the same structure handles whatever comes next: a new process, a new tool, the next big idea. Each one gets held to the same test, does it actually change the hours, and by how much, so the estimate keeps pace with how the work really gets done.
THE APP
Interested in seeing how Baseline works? Contact us and we’d love to give you demo.