01/05Entry point

WaterburyTalk to us
AdvertisingAd Doppler: we modeled a company’s ad spend in a changing market

PPC advertising bids move every day. We created a model that forecasts where the bid market is heading and an optimal strategy curve across multiple strategy preferences. Advertisers have to decide on tradeoffs between efficiency and scale; our client was launching a new product and wanted to be aggressive, but also had an efficiency baseline. Likely our client will move along the curve of the model to move towards a more efficient strategy. The model executes all strategies automatically through the advertising API.

ChessQueens: a chess improvement system

Queens helps people win at chess by giving a strategy for putting maximum practical pressure on their opponent based on the real data of similar skilled players. Current solutions largely focus on how to emulate “perfect play” from a computer, but humans play very differently. The person can study only the most consequential positions to cut down on study time enormously. This showcase has wider business application in problems where people need to prioritize a small amount of decisions from a large data set. Google also famously trained their systems on chess, as it is a rich and adversarial closed environment.

What the exit assumption is worth135 scenarios

exit \ growth0.51.01.52.02.53.03.54.04.5
4.7516.817.518.218.819.520.120.821.422.1
4.8816.116.717.418.018.719.320.020.621.3
5.0015.315.916.617.217.918.519.219.820.5
5.1314.415.115.716.417.017.718.319.019.6
5.2513.614.314.915.616.216.917.518.218.8
5.3812.813.414.114.715.416.016.717.318.0
5.5011.912.513.213.814.515.115.816.417.1
5.6311.011.712.313.013.614.314.915.616.2
5.7510.110.811.412.012.713.414.014.615.3
5.889.29.810.511.111.812.413.113.714.4
6.008.28.99.510.210.811.512.112.813.4
6.137.37.98.69.29.910.511.211.812.5
6.256.37.07.68.38.99.610.210.911.5
6.385.36.06.67.37.98.69.29.910.5
6.504.35.05.66.36.97.68.28.99.5

Coverage · noi over debt servicefloor 1.20x

coverage 1.25x against a 1.20x floorholds

subject to g(a) ≤ c

the drawn frontier is the preferred return

Distributions · the waterfallreconciles to 1.79x

y1y2y3y4y5exit
preferred · 8.0%1361571781992221,433
return of capital5,810
80 / 20 to 12%1,850
70 / 30 thereafter420
to the members892 across the hold9,513

The checks336 replayed · 3 wiring faults

  • every cell reproduced · replayed to the cent
  • three wiring faults found · one question, asked

illustrative figures · the live sheet replaces them

The same cash flows, quoted five waysone loan · five readings

unlevered8.2%
after debt service6.5%
levered12.7%
equity multiple1.79x
cash-on-cash · y12.3%

Amortization · 6.25% · thirty yearspmt 45.8

minterestprincipalbalance
01(38.75)7.067,432.9
03(38.68)7.137,418.7
05(38.60)7.217,404.3
07(38.53)7.287,389.8
09(38.45)7.367,375.1
11(38.37)7.447,360.3
13(38.30)7.517,345.3
15(38.22)7.597,330.2
17(38.14)7.677,314.9
19(38.06)7.757,299.4
21(37.98)7.837,283.8
23(37.90)7.917,268.0
25(37.81)8.007,252.0
27(37.73)8.087,235.9
29(37.64)8.167,219.6
31(37.56)8.257,203.2
33(37.47)8.347,186.5
35(37.39)8.427,169.7

principal rising a cent at a time · balance 6,983 at year five

Operating · month by month216 cells · $000s

price 12,400 · plan 850 · loan 7,440 · equity 5,810

mgrossvacopexnoidebtcash
01111.8(6.0)(49.7)56.1(45.8)10.4
02112.2(5.1)(48.6)58.5(45.8)12.8
03112.4(5.7)(49.2)57.5(45.8)11.7
04112.4(5.8)(49.9)56.7(45.8)10.9
05112.7(6.1)(49.9)56.7(45.8)10.9
06113.2(6.2)(50.2)56.8(45.8)11.0
07113.3(5.6)(53.2)54.5(45.8)8.8
08114.3(6.4)(51.4)56.4(45.8)10.7
09114.4(5.2)(50.4)58.8(45.8)13.1
10114.8(5.8)(50.9)58.1(45.8)12.4
11114.4(6.3)(50.4)57.7(45.8)11.9
12114.3(6.3)(50.9)57.2(45.8)11.4
y1noi 685debt 549cash 136
13115.2(5.8)(50.8)58.5(45.8)12.8
14115.2(6.2)(51.6)57.4(45.8)11.6
15116.4(6.0)(52.8)57.6(45.8)11.8
16116.6(5.5)(51.0)60.1(45.8)14.3
17115.9(5.7)(51.0)59.3(45.8)13.5
18116.2(6.1)(50.6)59.5(45.8)13.8
19117.4(5.9)(55.1)56.4(45.8)10.7
20117.6(6.2)(52.7)58.7(45.8)13.0
21117.3(6.4)(51.1)59.9(45.8)14.1
22117.5(5.7)(53.0)58.8(45.8)13.0
23118.1(5.9)(52.6)59.6(45.8)13.9
24117.7(6.3)(51.2)60.2(45.8)14.4
y2noi 706debt 549cash 157
25118.5(6.4)(52.1)60.0(45.8)14.3
26118.7(5.6)(52.5)60.6(45.8)14.9
27118.9(6.0)(53.2)59.7(45.8)14.0
28119.8(5.3)(54.4)60.1(45.8)14.4
··months 29–36 continue
y3noi 727debt 549cash 178

months 37–60 continue the same ledger · going-in 5.5% on price

Real estateUnderwriting alongside Soho Development

Soho Development needed to run a number of spreadsheets to evaluate investment opportunities across many years. Their old system was time intensive and could only feasibly produce a couple versions, a bull case and a bear case. Our model allowed them to navigate all of the possible scenarios to see which assumptions are load bearing for success. Soho could now, if they wanted, invest in properties other companies would pass on by understanding risk better.

We build custom artificial intelligence for your business.

We model one part of your business as mathematics, and the system gets better every time it runs.

For the technical reader
Niré BeautySlow MorningsQueensSoho Development

What just became possible

Contrary to popular belief, humans are great at predicting the weather.

As I sit writing this article in a cafe in Los Angeles, I can see dusty streaks on a window that hasn’t seen rain in months.

I’m not surprised; my weather app said it was likely to rain this morning.

As a fan of weather systems, I’m very excited to be able to accurately track an advertising budget as it passes day-by-day, similar to how a person might watch the colored radar of a storm move across the screen.

I call this project Ad Doppler, and I have the option to engage with it in multiple ways.

The easiest option is to do nothing and let the model execute the advertising strategy I co-signed earlier this month. The results are great.

But if I wanted, I could change strategy; perhaps I think it is time to be aggressive with our spend and not worry about efficiency so much. The Ad Doppler is already evaluating all of the optimal strategies and is ready to implement a new plan across all of the keywords immediately. It has already configured all of the bids and actions that would result in that action. The math that allows this is called a “surface”.

If I wanted, I could inspect all changes before they get pushed. In advertising, the implementation means setting bids for hundreds of keywords, which gets complicated. If I saw there was a keyword that has a very curious adjustment, something non-intuitive, I could ask the integrated LLM (eg: ChatGPT or Claude) to look at all of the context that came into this decision. Waterbury allows for better human LLM collaboration through this transparency.

However, I don’t want to change anything. The model is self-learning and executing on its own.

And the business results are following in an exciting way.... (unfortunately, I can’t share this clients data)

Other advertising AI solutions are trained on a lot of data, which can be useful, but that data is not your business and not your products. Newer AI solutions allow for fast and complicated “automations” that do a lot fast; but in our experience, these systems operate with a level of accuracy that we cannot cosign.

At Waterbury, we take an approach that we’re calling AI Architecture. We believe that it’s best to be disciplined, and especially if you’re paying for it, artificial intelligence should be intelligent, thoughtful, and get better with time, not just be fast and complicated.

We have noticed that even the phrase Artificial Intelligence has grown to mean “LLMs” (large language model) like ChatGpt and Claude. We use these same models and superpower them with a holistic approach that allows an LLM to be just the tip of the large intelligence iceberg underneath. As these models improve, our systems improve. We have found this to be very useful across a number of business problems. We think there are a number of places we’ve taken from mathematics and old engineering that current AI products undervalue. But you can leave the dusty textbook reading for us, we hope the value is obvious in the metrics you already review every day.

We want to partner with people who are excited as we are to see the limit of the systems we can create. Over time we want to continue to improve and create more systems that allow an advantage of competitors who are using “out of the box” AI products.

We hope to earn your trust over time and build together. We can only take on a few projects at a time, so we only take projects that are a good fit for both of us. We’d love to take a look at your business and see if it can see improvement.

Our clients seem to have a sense of pride that they have something that no other business has, something truly on the cutting edge and unique to them. That feeling has inspired me to continue this work.

I’m excited, and I hope you are too.

For the technical reader · positions we take

As much determinism as we can get. Anything with a right answer is computed by the infrastructure. LLMs should not be asked to multitask and improvise engineering solutions on the fly. Many problems are already solved problems in the context that they are intended. For example, deterministic data can easily be stored in a SQL database. Important decisions cannot rest on an infrastructure of LLMs playing telephone. A 90% success rate with 10 actions means a total success probability of 35%. At Waterbury we strive for as close to 100% as we can get. The LLM models in Waterbury never compute truth from scratch. At Waterbury, we find they are incredible at being well informed librarians who select, interpret, and recommend on the back of the infrastructure.

The model is the brain, and we build it a body. A frontier language model does reasoning, and we put it in a position to succeed. We have found we can send a multidimensional provenance packet that can be “unzipped” to inspect all of the mathematical components. Importantly, we can pass probabilities and other forms of math deterministically to accurately evaluate the amount of uncertainty in the real life business situation, without taking on probability of if the model will understand the problem accurately this time. In other words, for the problems we try to solve, we don’t hand over a pile of documents and ask the model: “no mistakes”

The packet carries any mathematics, including the uncertain kind. Estimates arrive as distributions with their spread and their assumptions. Probabilistic judgment is welcome here which is where the LLMs shine. But we care a lot about deterministic custody: what it was made from, when, by which data, and what happened after so it can improve with time.

The whole decision is mapped as a written chain of mathematics. Collaboration is much easier, accurate, and more effective when changing your assumptions automatically leads to a different conclusion without having to redo your work each time. We’ve had the experience of going in circles with an LLM as it keeps adjusting to your newest critique, while losing the precision you had a few chat sessions ago. We’ve found this “math spine” allows us a ton of visualization, replayability, and improvement possibilities that aren’t possible with an approach that either locks you into one way of operating, or requires luck that the logic is completely sound. We think an underrated part of intelligence is the ability to learn and improve, which out of the box LLMs cannot do.

Zoom into any term and it substitutes. We’ve gone through great lengths to make sure systems can be expressed as math. An important practical consideration is the ability to “zoom by substitution” which is the ability to have a complete unbroken logic chain that can get very complicated and have many inputs, without becoming an unintelligible mess that can’t be understood or improved.

The model’s working view may not quietly leave things out. We give the LLMs explicit consent to tell our clients when they’re not sure. We understand that LLMs many times cheat in impossible tasks and we want to have a friendly relationship with our LLM partners, allowing for honest communication both ways. Having as much of the system be deterministic allows for the models intelligence to be used on the problem itself, not on the infrastructure required to even pull up the problem into consideration.

Nothing durable lives inside any model. The configuration lives in the evidence, the definitions, and the record, on the bet that those appreciate as models improve while a fine-tune depreciates. Every prediction is stamped with the model that made it, so the system improving and a better model shipping never get confused for each other. As models improve, they understand our process and provide an even greater level of collaboration.

The finish line sits outside the model. A capable model under pressure starts optimizing for looking finished. So honest failure is a valid outcome, success is established from fresh evidence rather than the model’s claim to be done, and the watch reads pressure signals: repeated failures, drifting payloads, and widening authority. Errors and inefficiency are sent to Waterbury to be refined and improved.

We believe in verification We heavily use ideas from cryptography such as hashes to prove where information came from. We find that LLMs can sometimes lead you into “word soup” and it helps the LLMs from drifting when they can easily see the path information took.

The open problem: mapping mathematics to reality Many geniuses have created many types of math that are useful across many problems. We strive to employ the wisdom on the right kind of math describing just the kind of problem it was designed to describe. This is easier said than done, but this is where the majority of our value comes from compared to other frontier AI solutions.

Where this goes, stated as ambition. The research program behind the practice is machine epistemology: systems whose evidence, mathematics, interpretation, authority, action, outcome, and learning stay connected over time, and studied that way. Not how smart is the model, but what did the system observe, believe, predict, get authorized to do, actually do, and did the learning improve a later decision. The packet makes frontier progress additive: freeze the evidence and the task, swap model generations, measure the difference. Given a decade, the ladder runs from living laboratories to proof-carrying intelligence and verifiable institutions whose capabilities compound across organizations without collapsing anyone’s evidence or authority. None of that is product today, and we say so.

Let’s talk about your business

For inspiration

Some ideas of recurring problems we can solve for you

Advertising

Deploy your advertising budget daily

Replenishment

When do we reorder and how much

Assortment

What products in your niche are trending

Sales

Where in your sales calls do you lose customers

Underwriting

Which deals are worth the risk

Pricing

What is the profit optimal price

The question is written as mathematics. Its symbols become the parts of a structure: your records, the meanings agreed for them, the rules you set, and one answer. Any answer can then be walked back down to the records that produced it.= argmax𝔼[ Σ γᵏ|]subject tos′ ~ p( · | s, a, θ )a*v(s, a)Ig(a) ≤ cTHEN OBSERVE, AND SOLVE IT AGAIN TOMORROWYour recordsAgreed meaningsRules you setThe answer
Write the question. Watch it become the thing. Walk any answer back down.

The core

How we get results

Waterbury is an architecture practice for artificial intelligence, and the name is literal. We plan and build a math spine of how a business truly operates; the model hops in and becomes the brain. We start by writing your question (not the answer) down as mathematics.

This “math sketch” decides most of the value and is important to get right. We have found we are very strong in this critical moment.

We try to get the model off to a useful start by importing existing data. As the model gets smarter, to be able to follow what is happening. As important changes happen in the business, we may have to let the model know.

The system should respect limits, including your strategy. If you set a budget, it should respect and implement accordingly.

Reasoning comes from the strongest models available, the ones you already know: ChatGPT, Claude, Gemini, feel free to pick your favorite. The value we provide is everything that gets sent to that model. This can be an involved process, but it is only as involved as the problem needs to be optimally solved.

When you get an answer from a system we built, we want it to be honest. We want it to say when it doesn't know. And we want it to get better every day.

Make any wish

What do you wish your business could do?

Book a call

nick@waterbury.io