AI value or vanity? Why your AI isn’t delivering yet
Download the report
Request a DemoTry PiLog in

When More Data Does Not Automatically Mean a Better Prediction

Persis Duaik Tech Pre Sales
Publish date: 9th July 2026

How market value and fixture congestion led us to a more useful V3 confidence layer. 

60%
Overall accuracy stayed difficult
80%
Very high confidence tier accuracy
0
Live World Cup results used in training

Key takeaway 

V3 is not claiming football is suddenly easy to predict. It is saying something more useful: some predictions deserve more trust than others, and the analytics experience should show the difference. 

V3 is the story of what the data refused to prove 

Football prediction is a useful analytics problem because it is uncomfortable. There are obvious signals, there are tempting signals, and there are plenty of results that refuse to behave. That is exactly why we wanted the World Cup model to be visible inside Panintelligence: not just to show a prediction, but to show how the prediction changes when new ideas are tested. 

After V2, the strongest message was clear. The model still leaned heavily on the difference in FIFA points between two teams. That made sense, but it also raised a question: could we add something more current, more specific, or more human than a ranking score? 

The question behind V3 

V3 began with a simple ambition: improve the prediction by adding context that FIFA points might miss. FIFA points are useful because they summarize team strength over time. They are also broad, slow moving, and not very sensitive to some of the things football fans care about before a match. 

So, we tested two extra ideas inside Panintelligence, looking closely at how important each field became and how reliable the model was. The first was market value, as a proxy for player quality and squad depth. The second was fixture congestion, as a proxy for fatigue, rest, and recent match load. 

Experiment 1: market value 

  • The hypothesis: a team with more valuable players should have more talent, more depth, and a better chance of coping with tournament pressure. 
  • What we tested: squad value difference and a log-transformed version, using external Transfermarkt-style values from the Player Data file. The log version mattered because the richest squads can otherwise dominate the scale too aggressively. 
  • What happened: the feature was interesting context, but it did not add much independent prediction signal once FIFA strength was already in the model. 

Experiment 2: fixture congestion 

  • The hypothesis: teams with less rest or heavier recent match load may be more vulnerable, especially in tournament football. 
  • What we tested: rest days, short-rest pressure, and matches played in the last 14, 30, and 90 days. Current World Cup matches were excluded, so the live tournament remained a fair test. 
  • What happened: the idea made football sense, but the lift was tiny. Panintelligence did not show it as a meaningful driver of match outcome. 

The important result was not the result we wanted 

The tempting thing would have been to dress up one of these new fields as the V3 breakthrough. But if a feature does not earn its place in the model, the honest answer is not to force it. The honest answer is to learn from it. 

So V3 changed the question 

Instead of asking, "Can one more feature make the model dramatically more accurate?", V3 asks, "When this model makes a prediction, how much should we trust it?" 

That shift is small but powerful. Football has draws, upsets, injuries, red cards, and momentum swings. A single overall accuracy number hides too much. A confidence layer is more useful because it separates stronger predictions from dangerous ones. 

What confidence showed 

When historical predictions were grouped by confidence, the model became easier to interpret. The lowest confidence matches were genuinely risky. The highest confidence matches were much more dependable. 

Confidence Tier Accuracy What it Means
Very Low 41.1% Treat as high-risk. The model is not seeing enough separation.
Low 53.9% Useful, but needs caution and review before relying on it.
Medium 61.4% Close to the overall model level; helpful but not decisive.
High 65.6% Stronger evidence and a more dependable call.
Very High 80.0% The clearest tier: historically correct in four out of five matches.

Why this is still an upgrade 

At first, a model accuracy around 60% can feel disappointing. But for a three-way football outcome, that is not unusual. The practical problem is that the average hides the difference between a fragile prediction and a strong one. 

V3 improves the product because it gives the user a better decision surface. A very high confidence prediction can be treated differently from a low confidence prediction. One can support a clearer call. The other can trigger caution, review, or a different line of questioning. 

The analytics lesson 

The best part of this work is not that every idea improved the model. It is that every idea was testable. Market value and fixture congestion were both reasonable hypotheses. Panintelligence made it possible to test them, compare them, and decide not to overclaim them. 

That is what a good prediction journey should do. It should make the model more transparent, not more mysterious. It should show which assumptions survived contact with the data, and which ones became useful lessons instead. 

The V3 takeaway 

V3 helps users separate predictions worth trusting from predictions that deserve caution, debate, or a different model. That is a more honest upgrade than pretending one extra field solved football prediction. 

Topics in this post: 
Persis Duaik, Tech Pre Sales View all posts by Persis Duaik
Share this post
Related posts: 
Embedded Analytics

From Regulatory Burden to Growth Advantage: Rethinking Insurance Compliance

Picture this: the FCA and PRA want an auditable portfolio of regulatory reports covering pricing fairness, Consumer Duty outcomes, claims performance, solvency metrics and AI explainability.  You have 48 hours.  For most insurers, that single scenario is enough to trigger a scramble across underwriting systems, claims platforms, actuarial models and finance. Each system holds part of the truth. Each […]
Read more >>
Data visulization, Embedded Analytics

PiPredict: Why Simplicity Wins (and the Myth of the "Magical" Data Model)

As the 2026 FIFA World Cup heads into the final showdown on July 19 at MetLife Stadium, football fans are preparing for a heavyweight battle between Spain and Argentina. Spain swept past France 2-0, while Argentina defeated England 2-1.  But behind the scenes in our analytics lab, the real story is about how we tried to build a […]
Read more >>
Embedded Analytics, How to

How to Prepare Your Data for Predictive Analytics

One of the biggest misconceptions about predictive analytics is that the hard part is building the model. By the time you open any machine learning platform, most of the important work should already be done.  A good model starts with good data, and that means spending time understanding the business problem before thinking about algorithms. Whether you’re trying to […]
Read more >>
Houston... we've got mail.
Sign up with your email to receive news, updates and the latest blog articles to inspire you and your business.
© Panintelligence 2026