A benchmark is the fastest way to sense-check an estimate and the fastest way to fool a steering committee. The difference is whether anyone did the normalisation.
The seduction is the simplicity. One number, dollars per square metre or per megalitre per day, carrying the apparent authority of a built project. It fits on a slide, it travels by email, and it ends arguments. The trouble is that the number arrived stripped of everything that made it true: its year, its place, its scope and its story. Put those back and benchmarks become one of the most useful tools in estimating. Leave them off and the benchmark is deciding questions it does not understand.
What a benchmark is actually for
Benchmarks have two honest jobs. Early, at concept stage, they are a legitimate basis: a Class 5 estimate leans on the cost of analogous built work because there is nothing else to lean on, and that is what screening estimates are for: see A Class 5 estimate is a screening tool, not a budget. Later, they are validation: before an estimate is issued, it gets tested against comparable projects, which is part of the review discipline that AACE International Recommended Practice 31R-03, Reviewing, Validating, and Documenting the Estimate (RP 31R-03), describes. In both jobs, the benchmark’s value comes entirely from its basis being known.
What a benchmark is never for is replacing an estimate that measured actual quantities. When a bottom-up number and a benchmark disagree, the benchmark is a question, not a verdict.
The three adjustments
No comparison is allowed until three adjustments are made, in the open, with the arithmetic shown.
- Time: bring the comparison to the same pricing date with a relevant construction index, not general CPI. Construction inputs move on their own cycle, and over five years the gap between the right index and the consumer one is not a rounding error.
- Location: labour rates, productivity and logistics differ by more than most people expect between, say, Vancouver and Saint John, let alone between continents. A location factor with a source, applied deliberately.
- Scope: does the benchmark include owner’s costs, land, escalation and contingency? Half of the published ones do not say, which means half of the published ones are not yet usable numbers.
A worked example
An owner is validating a water treatment plant estimate: a 100 megalitre-per-day facility, estimated bottom-up at 310 million dollars. A well-documented comparable exists: a similar plant, built seven years ago in another province, at a recorded 195 million for 90 megalitres per day.
Normalise it. Time first: the relevant construction index has moved, call it, 30 percent over those seven years, taking the comparable to about 254 million. Location: the factor between the two places, from published data and labour agreements, call it 1.10, brings it to roughly 279. Scope: the recorded cost excluded the owner’s costs and land that the new estimate includes; adding the owner’s-cost share, call it twelve percent, lands the comparable near 313 million for 90 megalitres per day, or about 3.5 million per megalitre-day of capacity.
The estimate under review works out to 3.1. Two more comparables, normalised the same way, bracket the range at 2.9 to 3.6. The bottom-up number sits inside the normalised range, and the validation page says so with the arithmetic attached. Compare that with the version where someone divides 195 by 90 on a calculator and announces that the new plant is forty percent too expensive. Same data. Opposite conclusion. The difference was an hour of normalisation.
Then compare like with like
Even normalised, a benchmark only compares if the projects are the same kind of animal. A treatment plant benchmark in dollars per megalitre-day tells you very little if one plant is a greenfield build and the other an in-service upgrade with temporary works and night shifts. We keep the benchmarks that come with a scope description and discard the ones that come as a single number, however impressive their source. A benchmark you cannot interrogate is not data; it is a rumour with units.
A benchmark without its basis is an anecdote with a decimal point.
How we use them
As a range, never a point. Three to five normalised comparables define a band, and the bottom-up estimate is tested against the band. If it lands inside, good: the validation page records the comparables, the adjustments and the conclusion. If it lands outside, that is not a signal to move the estimate; it is a signal to find out which of the two is describing a different project. Sometimes the estimate is missing scope. Just as often, the comparables were simpler jobs than anyone remembered, and the estimate is the honest one. The investigation is the value; the band just tells you where to dig.
Where this goes wrong
The conference-slide number. A cost-per-unit figure from a presentation, origin unknown, scope unknown, quietly becomes the reference point for a region. It gets quoted for a decade because it is memorable, and memorable is not the same as true.
The CPI shortcut. Escalating a cost from several years ago to today with the consumer price index, because it was the index at hand. Construction indices exist, diverge from CPI for years at a time, and are no harder to look up.
The friendly comparable. Five comparables exist; the one that supports the desired conclusion gets the slide. Benchmarking that starts from the answer is advocacy, and reviewers can usually smell it.
Benchmarking against estimates. The “comparables” turn out to be other estimates, not built costs. Estimates validated against estimates inherit each other’s optimism, and the whole exercise measures agreement, not accuracy.
The sample of one. A single comparable, however good, is an anecdote. It takes a handful, normalised the same way, before the band means anything, and the basis of estimate should say how many stand behind it.
The hour that saves the meeting
Normalisation is not glamorous work: an index series, a location factor, a scope checklist, an hour per comparable. But it is the difference between a benchmark that informs a decision and one that ambushes it. The projects that get this wrong do not find out at the validation stage. They find out at tender, from the market, which charges considerably more than an hour.
Emerald Group maintains normalised benchmarks from delivered projects and validates estimates against them as an independent service for owners and engineering firms. If a number on a slide is about to steer your project, get in touch first.