Machine learning and Extreme Gradient Boosting

by Guest Contributor 3 min read October 24, 2018

This is an exciting time to work in big data analytics. Here at Experian, we have more than 2 petabytes of data in the United States alone. In the past few years, because of high data volume, more computing power and the availability of open-source code algorithms, my colleagues and I have watched excitedly as more and more companies are getting into machine learning. We’ve observed the growth of competition sites like Kaggle, open-source code sharing sites like GitHub and various machine learning (ML) data repositories.

We’ve noticed that on Kaggle, two algorithms win over and over at supervised learning competitions:

  • If the data is well-structured, teams that use Gradient Boosting Machines (GBM) seem to win.
  • For unstructured data, teams that use neural networks win pretty often.

Modeling is both an art and a science. Those winning teams tend to be good at what the machine learning people call feature generation and what we credit scoring people called attribute generation. We have nearly 1,000 expert data scientists in more than 12 countries, many of whom are experts in traditional consumer risk models — techniques such as linear regression, logistic regression, survival analysis, CART (classification and regression trees) and CHAID analysis. So naturally I’ve thought about how GBM could apply in our world.

Credit scoring is not quite like a machine learning contest. We have to be sure our decisions are fair and explainable and that any scoring algorithm will generalize to new customer populations and stay stable over time. Increasingly, clients are sending us their data to see what we could do with newer machine learning techniques. We combine their data with our bureau data and even third-party data, we use our world-class attributes and develop custom attributes, and we see what comes out. It’s fun — like getting paid to enter a Kaggle competition! For one financial institution, GBM armed with our patented attributes found a nearly 5 percent lift in KS when compared with traditional statistics.

At Experian, we use Extreme Gradient Boosting (XGBoost) implementation of GBM that, out of the box, has regularization features we use to prevent overfitting. But it’s missing some features that we and our clients count on in risk scoring. Our Experian DataLabs team worked with our Decision Analytics team to figure out how to make it work in the real world. We found answers for a couple of important issues:

  • Monotonicity — Risk managers count on the ability to impose what we call monotonicity. In application scoring, applications with better attribute values should score as lower risk than applications with worse values. For example, if consumer Adrienne has fewer delinquent accounts on her credit report than consumer Bill, all other things being equal, Adrienne’s machine learning score should indicate lower risk than Bill’s score.
  • Explainability — We were able to adapt a fairly standard “Adverse Action” methodology from logistic regression to work with GBM.

There has been enough enthusiasm around our results that we’ve just turned it into a standard benchmarking service. We help clients appreciate the potential for these new machine learning algorithms by evaluating them on their own data. Over time, the acceptance and use of machine learning techniques will become commonplace among model developers as well as internal validation groups and regulators.

Whether you’re a data scientist looking for a cool place to work or a risk manager who wants help evaluating the latest techniques, check out our weekly data science video chats and podcasts.

Related Posts

The Email Address as Your Most Powerful Identity Signal

The why behind Experian's acquisition of AtData What happens when a comprehensive email intelligence database joins a global leader in data, analytics and fraud prevention? The acquisition of AtData adds 25+ years of building a complete view of email as an identity signal. Financial institutions can recognize, engage and protect customers unlocking a new standard for the way their teams work and the customer experience. That's what Experian's acquisition of AtData delivers. How we got here Not all email addresses tell the same story. Some are newly created. Some exhibit bot-like patterns. Some are inconsistent with every other signal you have about that person. Imagine a real customer. You have a job. You shop online. You have a primary email from your employer, a personal Gmail you've used for 15 years, and an old Yahoo address you still use for shopping because you've been using it since college. You're an engaged customer who interacts with brands, makes purchases and pays bills on time. But each system sees a different version of you. When you apply for credit, the lender sees one email. When you shop, the retailer sees another. When you sign up for a service, you might use the third. For financial institutions: You slow down the approval process to manually verify identity or approve applicants without the full picture. For retailers: You can't tell which version of "customer" is the most engaged, so you either over-mail or under-serve. For fraud systems: Sees a new account created under one email and flags it as suspicious because it doesn't have the history. This was the original problem AtData was built to solve in 1999. Twenty-five years later, that problem didn’t go away, it became more complex. Email fragmentation and device sharing are more common, and identity theft is more sophisticated. Capabilities that now work together Experian has built sophisticated identity and fraud solutions backed by consumer data resources and decades of expertise in credit and risk. AtData brought the ability to assess whether an email address is trustworthy, reachable and consistent—at scale, in real time. Experian is now making email intelligence foundational, not optional. This matters for: Fraud prevention and risk management: Distinguishing a returning customer from a new threat. Knowing whether an email is newly created, exhibiting bot-like patterns or inconsistent with other identities is crucial. Compliance: Building audit trails that can explain identity decisions. Email data history and behavioral signals create the documentation needed to defend your decisions. Credit: Verifying identity in a world where traditional signals are shifting. Email signals provide a persistent, durable identifier that confirms who someone actually is. Marketing: Reaching the right person across email, mail and digital channels. Email intelligence reveals which addresses are actively engaged and reachable. Research shows email remains one of the highest-ROI marketing channels outperforming paid search and social advertising1. The problem every marketer faces: You end up burning budget on addresses that bounce, are unmonitored or are associated with users who never open mail. For credit marketing specifically, email enables faster, more targeted delivery of firm offers across channels, something that's increasingly important in a post-cookie world. "Email is a persistent identifier in a fragmented world. It's what connects a person's postal address, phones, devices, behaviors—the full picture of who they are. By embedding that into our infrastructure, we're not just adding another data point. We're fundamentally improving how businesses understand who their customers are."- Ashley Knight, Senior Vice President, Financial Services and Data Why now? AI is reshaping how decisions are made in every industry. Models are getting faster, more automated and more embedded in core workflows. But AI is only as effective as the data behind it. Fragmented data + fast models = faster, larger-scale misclassifications. In an era of synthetic identities, AI agents, deepfakes and AI-generated activity, the value of durable, persistent, real-world data signals has increased dramatically. Deloitte’s Center for Financial Services projects that generative AI could drive fraud losses in the U.S. up to $40 billion by 2027, a 32% growth rate since 2023. And email sits at the center of it with business email compromise already being one of the most common and costly fraud types. People change phones, move homes and swap devices, but they often hold onto their email for years. That's the signal that protects your business, and the one we've built into the core of how we help you make decisions with confidence. View the press release here

August 6, 2026 by Zohreen Ismail
Building Financial Opportunity Through Purpose-Driven Partnership

Discover how the National Urban League and Experian partner to expand financial literacy and create economic opportunity.

August 6, 2026 by Scarlet Nickel
2026 U.S. Identity and Fraud Report 

Explore key findings and insights from our newly released 2026 U.S. Identity and Fraud Report. Read more now!

August 5, 2026 by Laura Burrows

Subscribe to our Newsletter

Enter your name and email for the latest updates.

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Subscribe to our Newsletter

Don't miss out on the latest industry trends and insights!
Subscribe