No API? The screen is the API
A terminal from 1994 has no port, but it has a display. We read it, structure what it shows, and the model works from that.
Intelligence, plumbed in
Almost nobody fails at the model any more. They fail on four systems that disagree, a machine with no port, and no way to tell whether the answer was right. That is the work we do.
What actually has to exist
This is the order they have to be built in. It is also the order they get skipped in, and the order things fail in.
01 Something has to reach the data. A REST API if you are lucky, a nightly file if you are not, a screen if there is nothing else.
This is where most projects stall, six weeks in.
02 Four systems each call it something different and two of them disagree about what it means. One agreed meaning has to live somewhere.
Boring, unsellable, and the reason the answers come out right.
03 A warehouse, in your region, that you own. History matters: a model cannot learn your seasonality from last Tuesday.
Ordinary data engineering. No way around it.
04 The answer is usually in a note, a letter or a PDF. Finding the right three paragraphs is most of the work.
Chunking and ranking decide quality far more than the model does.
05 A process that runs on its own, on rails we lay: read, decide, act, log. Bounded, reversible, and it asks before anything irreversible.
Unbounded autonomy on your systems is not a feature.
06 A scored test set built from your own records, so accuracy is a number you can check rather than an impression from a demo.
Without this you have a demo, not a system.
The hard inputs
The awkward sources are not an exception in this work. They are most of it, and they are the reason a pilot that worked never shipped.
A terminal from 1994 has no port, but it has a display. We read it, structure what it shows, and the model works from that.
Inference on a box we specify and build, sealed, inside your network. No cloud call to make and nothing to call out to.
However hard, whatever it is
The first question is what is actually going wrong, and that takes sitting with you and reading everything. Six layers of plumbing are useless if the problem was diagnosed wrong.
01 Days where the work happens, not a workshop in a meeting room. We watch the job get done and write down the shortcuts nobody wrote down.
02 Your data, your rules, your vendors and their documentation, and the published research on your sector. We report what is actually in there.
03 Not which tool fixes this. What is actually causing it, taken apart until we reach the piece that cannot be divided further.
04 Weeks, not quarters. By this point we are not guessing what to build, and guessing is the thing that makes projects long.
Six things we will not do
Anyone can list capabilities. What a supplier declines to build tells you considerably more.
Plenty of the work we are asked for turns out to be a query and a screen. We say so. It costs us the bigger project and keeps the client.
Anything can be made to work on a hand-picked sample. We test against a spread of your real data, including the ugly rows, and we show you where it fails.
Reads are cheap to undo. Writes are not. Anything irreversible waits for a person, and the log says who approved it.
If a chat box is genuinely the right interface, fine. It usually is not. Most operational value lands as a screen, a queue or a nightly job nobody sees.
Models change every few months and the good ones are all adequate for most operational work. Your data is the variable that decides the outcome.
It stays in the region you choose. Where the rules require it, inference runs on your own hardware, air-gapped, with no call out at all.
How you would know it worked
Agreed before anything is built, measured on your own records, and reported per source rather than as one average that hides the bad one.
Straight answers
Nothing we say on a page. Ask us on the call what our eval set looks like, where the semantic layer lives, and what accuracy we measured per source. A team that has not done the work cannot answer those three.
Whichever fits, and it changes. That question matters less than people think for operational work: the good models are all adequate and your data decides the result. We will tell you when a smaller one on your own hardware is the better answer.
A scored test set built from your own records, agreed before we build. You get accuracy per source and per case type, not one flattering average, and you get the list of what it still gets wrong.
Yes. On-device or on-premise inference, air-gapped, on hardware we specify and can build. Nothing leaves the network. We have shipped against that constraint rather than argued with it.
Usually because the pilot proved the model and never touched the plumbing. Ask what the pilot did about the four systems that disagreed. If the answer is nothing, you learned about the model and not about your operation.
Not how we build it. We take out the retyping and the chasing, and the judgement stays where it is. A system that removes the person who understood the work is a system that fails the year after.
The connector, the warehouse, the screen the team actually uses, and the sensor when nothing is measuring the thing you want to predict.
See everything we build
Software
Hardware
Ways of working
Whole ventures