Available as AI Engineer & AI Solutions Architect

Back to home

Blog

Everyone Wants an AI Prediction. Who Checks the Data Behind It?

A price prediction looks simple from the outside. Give AI some past sales, describe a product and ask what it should cost.

But where did those numbers come from? Are we comparing similar things? And what happens when the data is wrong?

Recently, I was working on a project where I was given a dataset of cars and the prices they were sold for. The task was to build a price prediction system from those records.

A dealer could use something like this when pricing stock. A buyer or seller could use it as a starting point for a negotiation.

The idea sounded straightforward: describe a car, find similar examples and estimate a price.

Car price prediction system

What happens behind the scenes?

I built the system on AWS, using S3 to store the data, Lambda to handle the workflow and SageMaker to train the model.

The process starts when we upload a data of a cars and prices they were sold for. That upload triggers a Lambda function that normalizes data. Next, another Lambda function starts a SageMaker training job to build a new nearest-neighbor model.

During training, the system selects settings such as feature weights, the number of comparable cars to use and how strongly closer matches should influence the price.

The main job of the model is to find similar cars and use them to predict the price.

What counts as a similar car?

I used a nearest-neighbor algorithm.

It finds cars from the same brand, compares the details the user supplied and uses the closest matches to calculate an estimate.

The closer the match, the more influence it has on the predicted price.

But not every detail matters equally. In this model, horsepower has more weight than color. Two cars should not become strong matches just because both happen to be blue.

So the task was more than averaging some prices. I needed to define what “similar” meant and how much each match should count.

And a good matching method still cannot fix a wrong price in the source data.

Show where the number came from

I wanted someone using the system to be able to inspect the result.

Under the estimate, the website explains which details were used and why those records were matched.

Prediction explanation showing the details used to match comparable cars

The user can also open the comparable cars and see the recorded prices behind the prediction.

Comparable cars and their recorded prices

This gives them something concrete to check. Do these comparisons make sense? Is an important detail missing?

When there are too few comparable records, the system says there is not enough data. Producing a number anyway would not make it more useful.

Check where it gets things wrong

I used a Jupyter Notebook to validate the model’s accuracy by comparing predicted prices with the actual recorded prices from test data.

This helped me understand how close the predictions were to the real values.

Model validation comparing predicted and recorded car prices

A prediction is not enough

A number without context is hard to trust.

If AI gives a prediction, show the user why: what data it used, what it compared and what influenced the result.

This gives people something they can check instead of asking them to trust the model blindly.

Don’t just show the answer. Show enough evidence to understand it.


I’m Dima - AI Engineer & Solutions Architect. I help teams turn complex processes into one working system. You can see more of my work at mikheiev.com