Machine learning and Extreme Gradient Boosting

by Guest Contributor 3 min read October 24, 2018

This is an exciting time to work in big data analytics. Here at Experian, we have more than 2 petabytes of data in the United States alone. In the past few years, because of high data volume, more computing power and the availability of open-source code algorithms, my colleagues and I have watched excitedly as more and more companies are getting into machine learning. We’ve observed the growth of competition sites like Kaggle, open-source code sharing sites like GitHub and various machine learning (ML) data repositories.

We’ve noticed that on Kaggle, two algorithms win over and over at supervised learning competitions:

  • If the data is well-structured, teams that use Gradient Boosting Machines (GBM) seem to win.
  • For unstructured data, teams that use neural networks win pretty often.

Modeling is both an art and a science. Those winning teams tend to be good at what the machine learning people call feature generation and what we credit scoring people called attribute generation. We have nearly 1,000 expert data scientists in more than 12 countries, many of whom are experts in traditional consumer risk models — techniques such as linear regression, logistic regression, survival analysis, CART (classification and regression trees) and CHAID analysis. So naturally I’ve thought about how GBM could apply in our world.

Credit scoring is not quite like a machine learning contest. We have to be sure our decisions are fair and explainable and that any scoring algorithm will generalize to new customer populations and stay stable over time. Increasingly, clients are sending us their data to see what we could do with newer machine learning techniques. We combine their data with our bureau data and even third-party data, we use our world-class attributes and develop custom attributes, and we see what comes out. It’s fun — like getting paid to enter a Kaggle competition! For one financial institution, GBM armed with our patented attributes found a nearly 5 percent lift in KS when compared with traditional statistics.

At Experian, we use Extreme Gradient Boosting (XGBoost) implementation of GBM that, out of the box, has regularization features we use to prevent overfitting. But it’s missing some features that we and our clients count on in risk scoring. Our Experian DataLabs team worked with our Decision Analytics team to figure out how to make it work in the real world. We found answers for a couple of important issues:

  • Monotonicity — Risk managers count on the ability to impose what we call monotonicity. In application scoring, applications with better attribute values should score as lower risk than applications with worse values. For example, if consumer Adrienne has fewer delinquent accounts on her credit report than consumer Bill, all other things being equal, Adrienne’s machine learning score should indicate lower risk than Bill’s score.
  • Explainability — We were able to adapt a fairly standard “Adverse Action” methodology from logistic regression to work with GBM.

There has been enough enthusiasm around our results that we’ve just turned it into a standard benchmarking service. We help clients appreciate the potential for these new machine learning algorithms by evaluating them on their own data. Over time, the acceptance and use of machine learning techniques will become commonplace among model developers as well as internal validation groups and regulators.

Whether you’re a data scientist looking for a cool place to work or a risk manager who wants help evaluating the latest techniques, check out our weekly data science video chats and podcasts.

Related Posts

Are Fraudsters Building Better Identities Than Your Customers?

Fraudsters are getting surprisingly good at onboarding. Sometimes, better than your customers. Legitimate customers treat onboarding like an errand. They start an application between other tasks, get distracted, forget a password, switch devices, upload a document or come back later to finish. Their digital lives aren’t always linear, because real life isn’t either. Fraudsters approach onboarding differently. For them, opening an account is the objective. Every interaction is designed to increase the odds of success. The difference raises an uncomfortable question hanging over onboarding: What exactly are we rewarding? When smooth becomes suspicious Digital onboarding has traditionally rewarded experiences that feel smooth, consistent and complete. The challenge is that legitimate customers rarely behave that way. Most people approach onboarding somewhere between mildly distracted and mildly annoyed. They pause halfway through because dinner is burning. They reopen an old account only to realize everything is attached to an email they made in college and, somehow, still use for airline receipts. Digital life accumulates history unevenly, because ordinary life does too. Fraudsters have every reason to eliminate those inconsistencies. Applications may be rehearsed. Identity attributes are assembled deliberately. Contact points are prepared in advance. Every interaction is optimized to make the application appear credible. Ironically, the qualities organizations often associate with confidence — clean submissions, steady progression and few corrections — can also describe applications that have been carefully engineered to pass inspection. The challenge isn't that smooth onboarding is meaningless. It's that smooth onboarding, by itself, doesn't tell the whole story. Context changes interpretation A smooth onboarding experience should be the beginning of the evaluation, not the end. Behavior provides important context. How someone moves through an application can reveal whether the experience feels naturally human or unusually orchestrated. Do they interact naturally? Do they hesitate, correct mistakes or navigate in ways that resemble ordinary human behavior? Or does the session appear unusually scripted, automated or repetitive? Identity verification adds another layer. Matching information across trusted sources, validating identity details and strengthening confidence in account creation remain important, particularly when onboarding decisions carry financial, fraud or customer experience consequences. But verification largely answers a point-in-time question: Does this information match right now? A third layer comes from digital history. An inbox attached to years of airline receipts, loyalty accounts, subscription renewals, account recovery, financial notifications and familiar digital routines introduces a different kind of confidence. Legitimate digital identities leave behind patterns of persistence and engagement that develop gradually over time. Fraudsters can assemble convincing identity attributes, but creating years of ordinary digital life is much harder. Building confidence in an identity requires more than verifying information submitted during a single onboarding session. It requires understanding whether the identity reflects a broader history that supports what the application suggests. A multilayered approach builds stronger identity confidence No single signal can provide a complete view of identity risk. Organizations need multiple sources of confidence that reinforce one another. That's the thinking behind our approach: combining behavioral intelligence, identity verification and digital identity continuity into a more complete view of risk. We bring these complementary layers together through: • NeuroID adds behavioral context during onboarding and account creation, helping identify interaction patterns that may indicate automation, manipulation or coordinated fraud. • Precise ID® strengthens identity verification and resolution by comparing applicant information with trusted identity data. • AtData, recently added to our portfolio, contributes email-centered intelligence based on persistence, engagement and long-term digital history. Together, these capabilities help organizations move beyond evaluating a single moment in time to understanding whether an identity is supported by consistent behavior, trusted identity data and an established digital history. The future of fraud prevention isn't about rewarding the smoothest application. It's about recognizing the most trustworthy identity. Fraudsters can rehearse an application. They can optimize an onboarding journey. They can even assemble convincing identity attributes. What they can't easily manufacture is years of ordinary digital life. That's why digital identity continuity has become an important layer of modern fraud prevention. Combined with identity verification and behavioral intelligence, it helps organizations distinguish between identities that simply look convincing and those supported by a history that is much harder to fake. Learn more Contact us

September 2, 2026 by Julie Lee
From Hybrids to Refinancing: Consumers are Finding New Roads to Vehicle Affordability

For today’s automotive consumers, considering a vehicle purchase isn’t just about the price they see on the window, it’s about finding the right combination of their vehicle preference and monthly payment. In fact, data from Experian Automotive’s State of the Automotive Finance Market Report: Q2 2026 highlighted how affordability continues to shape the automotive finance market. For instance, hybrids offered the lowest average new vehicle loan payment across all fuel types, coming in at $646 in Q2 2026, compared to electric vehicles (EVs) at $692, and gasoline-powered vehicles at $721. This led to considerable growth in new vehicle market share for hybrids this quarter, accounting for 16.80%, from 12.99% last year. While the automotive market continues to offer consumers an expanding mix of fuel types, the combination of growing hybrid share and comparatively lower monthly payments is something worth watching. Affordability isn’t just about what consumers drive, it’s how they finance it While hybrid vehicles are continuing to pave their way in the vehicle market, consumers who already have an auto loan are finding greater savings through refinancing. In the second quarter of 2026, automotive refinancing reached approximately 140,000 loans. More notably, the financial benefit associated with refinancing has grown. Consumers who refinanced this quarter reduced their average interest rate by more than 2.4%, with the average rate moving from 10.40% on the original loan to 7.97% on the refinanced loan. Those rate reductions translated into meaningful monthly savings, especially when refinancing through particular lenders. In Q2 2026, refinancing saved consumers an average of $83 per month, compared to an average monthly savings of $64 this time last year. However, credit unions delivered the largest average payment difference among lender types at $102 this quarter, followed by banks ($65), and finance companies ($38). It’s important for automotive professionals to acknowledge that affordability is not a single moment in the vehicle journey. It can influence the vehicle a consumer chooses, the financing they opt for during that transaction, and the decisions they make years after driving off the lot. Understanding and leveraging those different moments can help professionals identify opportunities to better serve consumers throughout the vehicle ownership lifecycle. To learn more about automotive finance trends, view the full State of the Automotive Finance Market Report: Q2 2026 presentation on demand.

August 27, 2026 by Melinda Zabritski
AI Agent Identity Verification: How to Verify AI Agents in Digital Transactions

AI agents are changing the way consumers interact with businesses online. Learn how you can establish greater confidence in AI transactions.

August 26, 2026 by Laura Burrows

Subscribe to our Newsletter

Enter your name and email for the latest updates.

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Subscribe to our Newsletter

Don't miss out on the latest industry trends and insights!
Subscribe