Data & Analytics
By: Barbara Rivera Every day, 2.5 quintillion bytes of data are created – in fact, 90% of the world’s data was created in only the last few years. With the staggering amount of data available, we have an unprecedented opportunity to uncover new insights and improve the way our world functions. The implications of these new capabilities are perhaps nowhere else as crucial as within our government. Public sector officials carry the great responsibility of conducting complex missions that directly affect our communities, our economy, and our nation’s future. The ability to make more informed, insightful choices and better decisions is paramount. Especially at a time of broader global unrest and uncertainty, Americans rely on our government to be transparent, fair, ready and to make the right decisions – our trust is in the hands of our elected officials and public servants. Data alone is not enough to inform and affect change. However, with integrated information assets, insightful analysts and collaborative processes, data can be transformed into something meaningful and actionable. Our government has already begun leveraging data for good across agencies and varied missions, with more potential unlocked each day. Local governments like Orange County, California are utilizing data through address verification services to keep their voting lists accurate – ensuring the integrity of elections and saving the taxpayers thousands of dollars otherwise wasted on mailings to outdated lists. The Orange County Registrar of Voters – the fifth largest voting jurisdiction in the county – has been able to cancel 40,000 voting records, with an estimated savings of $94,000 expected from 2012 through 2016. The examples are numerous and growing: A suite of optimization tools helps states find non-custodial parents, determine their capacity and likelihood to pay child support, and trigger alerts with new critical information, maximizing the likelihood of payment and recovery, ultimately improving the welfare of children and reducing poverty More than 150 state, county and local law enforcement agencies leverage data to help identify persons of interest, conduct background screening for employees and contractors and provide financial backgrounds for criminal investigations, ensuring our continued safety By using the power of data to manage user authentication, credentials and access controls, the government is working harder – and smarter – to protect our security The government is leveraging verified commercial data to help agencies validate the fiscal responsibility of potential contractors and monitor existing contractors, which helps provide transparency and reduce risk By using data and analytics to authenticate applicants and validate financial data, the government is ensuring access to benefits for those who meet eligibility requirements, while at the same time reducing fraud Private sector partners are supporting municipal efforts to improve financial stability in households by providing the current credit standing of consumers and monitoring overall changes in financial behaviors over time, to help counsel and educate citizens And that’s only the beginning. The possibilities are endless – from healthcare to finance to energy – data can be leveraged for the advancement of our society. It even happens behind the scenes, working to protect information in ways most citizens never realize. Data insights are used to ensure citizens have secure online access to their information – ever see those randomized, personal questions? That’s data at work. The same technology is the de facto ID Proofing standard for the VA and CMS. How does it all work? By combing through the data carefully, putting it in context, looking at it in new ways, and thinking about what all this information really means. Much of this is made possible through public-private partnerships between the government and companies like Experian. So the next time someone complains about the slow pace of government, let them know the truth is government is moving quickly, leveraging data and private sector partnerships to uncover new insights that impact the greater good.
This is the third post in a three-part series. Experian® is not a doctor. We don’t even play one on TV. However, because of our unique business model and experience with a large number of data providers, we do know data governance. It is a part of our corporate DNA. Our experiences across our many client relationships give us unique insight into client needs and appropriate best practices. Note the qualifier — appropriate. Just as every patient is different in his or her genetic predispositions and lifestyle influences, every institution is somewhat unique and does not have a similar business model or history. Nor does every institution have the same issues with data governance. Some institutions have stabile growth in a defined footprint and a history of conservative audit procedures. Others have grown quickly through aggressive acquisition marketing plans and unique channels and via institution acquisition/merger, leading to multiple receivable systems and data acquisition and retention platforms. Experian has provided valuable services to both environments many times throughout the years. As the regulatory landscape has evolved, lenders/service providers demand a higher level of hands-on experience and regulatory-facing credibility. Most recently, lenders have required assistance on the issues driven by mandates coming from the Comprehensive Capital Analysis and Review (CCAR), Office of the Comptroller of the Currency (OCC) and the Consumer Financial Protection Bureau (CFPB) bulletins and guidelines. Lenders are best served to begin their internal review of their data governance controls with a detailed individual attribute audit and documentation of findings. We have seen these reviews covering fewer than 200 attributes to as many as more than 1,000 attributes. Again, the lender/provider size, analytic sophistication and legacy growth and IT issues will influence this scope. The source and definition of the attribute and any calculation routines should be fully documented. The life cycle stage of attribute acquisition and usage also is identified, and the fair lending implication regarding the use of the attribute across the life cycle needs to be considered and documented. As part of this comprehensive documentation, variances in intended definition and subsequent design and deployment are to be identified and corrective action guidance must be considered and documented for follow-up. Simultaneously, an assessment of the current risk governance policies, processes and documentation typically is undertaken. A third party frequently is leveraged in this review to ensure an objective perspective is maintained. This initiative usually is a series of exploratory reviews and a process and procedures assessment with the appropriate management team, risk teams, attribute design and development personnel, and finally business and end-user teams, as necessary. From these interviews and the review of available attribute-level documentation, documents depicting findings and best practices gap analysis are produced to clarify the findings and provide a hierarchy of need to guide the organization’s next steps: A more recent evolution in this data integrity ecosystem is the implication of leveraging a third party to house and manipulate data within client specifications. When data is collected or processed in “the cloud,” consistent data definitions are needed to maintain data integrity and to limit operational costs related to data cleansing and cloud resource consumption. Maintaining the quality of customer personal data is a critical compliance and privacy principle. Another challenge is that of maintaining cloud-stored data in synchronization with on-premises copies of the same data. Delegation to a third party does not discharge the organization from managing risk and compliance or from having to prove compliance to the appropriate authorities. In summary, a lender/service provider must ensure it has developed a rigorous data governance ecosystem for all internal and external processes supporting data acquisition, retention, manipulation and utilization: A secure infrastructure includes both physical and system-level access and control. Systemic audit and reporting are a must for basic compliance standards. If data becomes corrupted, alternative storage, backup or other mechanisms should be available to protect the information. Comprehensive documentation must be developed to reveal the event, the causes and the corrective actions. Data persistence may have multiple meanings. It is imperative that the institution documents the data definition. Changes to the data must be documented and frequently will lead to the creation of a new data attribute meeting the newer definition to ensure that usage in models and analytics is communicated clearly. Issues of data persistence also include making backups and maintaining multiple archive copies. Periodic audits must validate that data and usage conform to relevant laws, regulations, standards and industry best practices. Full audit details, files used and reports generated must be maintained for inspection. Periodic reporting of audit results up to the board level is recommended. Documentation of action plans and follow-up results is necessary to disclose implementation of adequate controls. In the event of lost or stolen data, appropriate response plans and escalation paths should be in place for critical incidents. Throughout this blog series, we have discussed the issues of risk and benefits from an institution’s data governance ecosystem. The external demands show no sign of abating. The regulators are not looking for areas to reduce their oversight. The institutional benefits of an effective data governance program are significant. Discover how a proven partner with rich experience in data governance, such as Experian, can provide the support your company needs to ensure a rigorous data governance ecosystem. Do more than comply. Succeed with an effective data governance program.
Data quality continues to be a challenge for many organizations as they look to improve efficiency and customer interaction.
As part of its guidance, the Office of the Comptroller of the Currency recommends that lenders perform regular validations of their credit score models in order to assess model performance.
Using a risk model based on older data can result in reduced predictive power.
By: Maria Moynihan Crime prevention and awareness techniques are changing and data, analytics and use of technology is making a difference. While law enforcement departments continue to face issues related to data - ranging from working with outdated information, inability to share data across departments, and difficulty in collapsing data for analysis - a new trend is emerging where agencies are leveraging outside data sources and analytic expertise to better report on crimes, collapse information, predict patterns of behavior and ultimately locate criminals. One best practice being implemented by law enforcement agencies is to skip trace an individual much like a debt collector would. Techniques involve using historic address information and individual connections to better track to a person’s current location. See the full write up from CollectionsandCreditRisk.com to see how this works. Another great example of effective use of data in investigations can be seen in this video, where one Experian client, Intellaegis of El Dorado Hills, CA, recently worked with local law enforcement to follow the digital data footprints of a particular suspect, finding her in in just five minutes of searching. p> And, yet another representation of improved data gathering, handling and sharing of information for crime prevention and awareness can be found on a site I was just made aware of by one of my neighbors - www.crimemapping.com. Information is collapsed across departments for greater insight into the crimes that are happening within a neighborhood, offering a more comprehensive option for the general public to turn to on local area crime activity. Clearly, data, analytics and technology are making a positive impact to law enforcement processes and investigations. What is your public safety organization doing to evolve and better protect and serve the public?
Data quality should be a priority for retailers at any time of the year, but even more so as the holiday season approaches. According to recent research from Experian, organizations feel that, on average, 25 percent of their data is inaccurate and 12 percent of departmental budgets are wasted due to inaccuracies in contact data. During the 2013 holiday season, consumer spending is expected to increase by at least 11 percent. Retailers need to improve data quality early on in order to ensure that relevant holiday offers reach consumers and to take advantage of the expected increase in consumer spending. View our recent Webinar: Unique insights on consumer credit trends and the impact of consumer behavior on the economic recovery Source: View our data quality infographic: ’Twas the month before the holidays
The average bankcard balance per consumer in Q2 2013 was $3,831, a 1.3 percent decline from the previous year. Consumers in the VantageScore® near prime and subprime credit tiers carried the largest average bankcard balances at $5,883 and $5,903 respectively. The super prime tier carried the smallest average balance at $1,881.
Using data from IntelliViewSM, Credit.com recently compiled a list of states with the highest average bankcard utilization rates. Alaska took first place, with an average utilization ratio of 27.73 percent. This should come as no surprise since Alaska has recently topped lists for highest credit card balances and highest revolving debt.
When validating a model in the presence of overlay criteria, it is important to remember that any metrics computed at the aggregate portfolio level will not be indicative of the model's true performance. While traditional validation methodologies and portfolio metrics may provide directional insight into model performance, the overlay strategy is an additional variable that must be accounted for in each step of the validation analysis. An effective validation should include: Establishment of an appropriate base line Piece-wise validation of overlay segments An overlay strategy analysis Do you have model validation questions? Learn more and transform your business goals with Experian's Analytical Consulting Services. Source: VantageScore® Solutions LLC white paper: Validating a Credit Score Model in Conjunction with Additional Underwriting Criteria. VantageScore® is owned by VantageScore Solutions, LLC.
As part of its expanded guidance, the Office of the Comptroller of the Currency explicitly recommends that financial services firms utilizing predictive models and decision analytics run regular validations to gauge model efficacy. The VantageScore® credit score model was recently measured against the best credit score models from each of the three largest credit reporting companies (CRCs). When comparing KS values, there is exceptionally strong performance for mortgage originations, with the VantageScore® credit score model outperforming the CRC models in a range from 8 percent to 12 percent. The average range of outperformance is 3 percent to 4 percent across the board for most of the key industries. View the VantageScore® Webinar: Executing Effective Validations in 2011 and Beyond. Source: Executing Effective Validations, American Banker. VantageScore® is owned by VantageScore Solutions, LLC.
VantageScore® Solutions LLC polled risk professionals about how they are measuring score performance, and 60 percent of respondents said they are now using metrics beyond the Kolmogorov-Smirnov (KS) statistic value. One new metric is score consistency, which is defined as the ability to provide near-identical risk assessment of a consumer across multiple credit reporting agencies. In other words, this means having confidence that when a consumer gets a 700 from one agency, he or she is likely to get a 700 from another agency. The other metric that risk managers referenced was stability, which is defined as the ability of a model to retain its predictive accuracy across an extended time frame. Learn more about the VantageScore credit score® Source: VantageScore newsletter, April 2011 VantageScore® is owned by VantageScore Solutions, LLC
Experian® QAS®, a leading provider of address verification software and services, recently released a new benchmark report on the data quality practices of top online retailers. The report revealed that 72 percent of the top 100 retailers are using some form of address verification during online checkout. This third annual benchmark report enables retailers to compare their online verification practices to those of industry leaders and provides tips for accurately capturing email addresses, a continuously growing data point for retailers. To find out how online retailers are utilizing contact data verification, download the complimentary report 2012 Address Verification Benchmark Report: The Top 100 Online Retailers. Source: Press release: Experian QAS Study Reveals Prevalence of Real-Time Address Verification Increasing Among Top Online Retailers.
By: John Straka For many purposes, national home-price averages, MSA figures, or even zip code data cannot adequately gauge local housing markets. The higher the level of the aggregate, the less it reflects the true variety and constant change in prices and conditions across local neighborhood home markets. Financial institutions, investors, and regulators that seek out and learn how to use local housing market data will generally be much closer to true housing markets. When houses are not good substitutes from the viewpoint of most market participants, they are not part of the same housing market. Different sizes and types and ages of homes, for example, may be in the same county, zip code, block, or even right next door to each other, but they are generally not in the same housing market when they are not good substitutes. This highlights the importance of starting with detailed granular information on local-neighborhood home markets and homes. To be sure, greater granularity in neighborhood home-market evaluation requires analysts and modelers to deal with much more data on literally hundreds of thousands of neighborhoods in the U.S. It is fair to ask if zip-code level data, for example, might not be generally sufficient. Most housing analysts and portfolio modelers, in fact, have traditionally assumed this, believing that reasonable insights can be gleaned from zip code, county-level, or even MSA data. But this is fully adequate, strictly speaking, only if neighborhood home markets and outcomes are homogenous—at least reasonably so—within the level of aggregation used. Unfortunately, even at zip-code level, the data suggests otherwise. Examples All of the home-price and home-valuation data for this report was supplied by Collateral Analytics. I have focused on zip7s, i.e. zip+2s, which are a more granular neighborhood measure than zip codes. A Hodrick-Prescott (H-P) Filter was applied by Collateral Analytics to the raw home-price data in order to attenuate short-term variation and isolate the six-year trends. But as we’ll see this dampening still leaves an unrealistically high range of variation within zip codes, for reasons discussed below. Fortunately there is an easy way to control for this, which we’ll apply for final estimates of the range of within-zip variation in home-price outcomes. The three charts below show the H-P filtered 2005-2011 percent changes in home-price per square foot of living area within three different types of zip codes in San Diego county. Within the first type of zip code, 92319 in this case, the home-price changes in recent years have been relatively homogenous, with a range of -56% to -40% home-price change across the zip7s (i.e., zip+2s) in 92319. But the second type of zip code, illustrated by 92078, is more typical. In this type of case the home-price changes across the zip7s have varied much more. The 2055-2011 zip7 %chg in home prices within 92078 have varied by over 40 percentage points, from -51% to -10%. In the third type of zip code, less frequent but surprisingly common, the home-price changes across the zip7s have had a truly remarkable range of variation. This is illustrated here by zip code 92024 in which the home price outcomes have varied from -51% to +21%, or a 71 percentage point range of difference—and this is not the zip code with the maximum range of variation observed! All of the San Diego County zip codes are summarized in the bar chart below. Nearly two-thirds of the zip codes, 65%, have more than 30 percentage points within-zip difference in the 2005-2011 zip7 %changes in home prices. 40% have more than a 40 percentage point range of different home-price outcomes, 23% have more than a 50 percentage point range, and 13% have more than a 70 percentage point range of differences. The average range of the zip7 within-zip code differences is a 37 percentage point median, 41 percentage-point mean. These high numbers are surprising, and are most likely unrealistically high. Summary of Within-Zip (Zip+2 level) Ranges of Variation in Home-Price Changes in San Diego: Percentage of Zips by Range Across Zip+2s in Home Price/Living Area %Change 2005-2011 Controlling for Factors Inflating the Range of Variation Such sizable differences within a typical single zip code clearly suggest materially different neighborhood home markets. While this qualitative conclusion is supported further below, the magnitudes of the within-zip variation in home-price changes shown above are quite likely inflated. There is a tendency for a limited number of observations in various zip7s to create statistical “noise” outliers, and the inclusion of distressed property sales here can create further outliers, with cases of both limited observations and distress sales particularly capable of creating more negative outliers that are not representative of the true price changes for most homes and their true range of variation within zip codes. (My earlier blog on June 29th discussed the biases from including distressed property sales while trying to gauge general price trends for most properties.) Fortunately, I’ve been able to access a very convenient way to control for these factors by using the zip7 averages of Collateral Analytics’ AVM (Automated Valuation Model) values rather than simply the home price data summarized above. These industry-leading AVM home valuations have been designed, in part, to filter out statistical noise problems. The bar chart below shows the still significant zip7 ranges within San Diego County zip codes using the AVM values, but the distribution is now shifted considerably, and more realistically, to a much smaller share of the zip codes with remarkably high zip7 variation. Compared with the chart above, now just 1% of the zips have a zip7 range greater than 60 percentage points, 5% greater than 50, and 11% greater than 40, but there are still 36% greater than 30. To be sure, this distribution, and the average range of zip7 differences—which is now a 25 percentage-point median, 26 percent age-point mean—do show a considerable range of local home market variation within zip codes. It seems fair to conclude that the typical zip code does not contain the uniformity in home price outcomes that most housing analysts and modelers have tended to simply assume. The difference between the effects on consumer wealth and behavior of a 10% home price decline, for example, vs. a 35 to 50% decline, would seem to be sizable in most cases. This kind of difference within a zip code is not at all unusual in these data. How About a Different Type of Urban Area—More Uniform? It might be thought that the diversity of topography, etc., across San Diego County (from the sea to the mountains) makes its variation of home market outcomes within zip codes unusually high. To take a quick gauge of this hypothesis, let’s look at a more topographically uniform urban area: Columbus, Ohio. When I informally polled some of my colleagues asking what their prior belief would be about the within-zip code variation in home price outcomes in Columbus vs. San Diego County, there was unanimous agreement with my prior belief. We all expected greater within-zip uniformity in Columbus. I find it interesting to report here that we were wrong. Both the H-P filtered raw home-price information and the AVM values from Collateral Analytics show relatively greater zip7 variation within Columbus (Franklin County) zip codes than in San Diego County. The bar chart below shows the best-filtered, most attenuated results, the AVM values. 5% of the Columbus zips have a zip7 range greater than 70 percentage points, 8% greater than 60, 23% greater than 50, 35% greater than 40, and 65% greater than 30. The average range of zip7 within-zip code differences in Columbus is a 35 percentage point median, 38 percentage-point mean. Conclusion These data seem consistent with what experienced appraisers and real estate agents have been trying to tell economists and other housing analysts, investors, and financial institutions and policymakers for quite a long time. Although they have quite reasonable uses for aggregate time-series and forecasting purposes, more aggregate-data based models of housing markets actually miss a lot of the very real and material variation in local neighborhood housing markets. For home valuation and many other purposes, even models that use data which gets down to the zip code level of aggregation—which most analysts have assumed to be sufficiently disaggregated—are not really good enough. These models are not as good as they can or should be. These facts are indicative of the greater challenge to properly define local housing markets empirically, in such a way that better data, models, and analytics can be more rapidly developed and deployed for greater profitability, and for sooner and more sustainable housing market recoveries. I thank Michael Sklarz for providing the data for this report and for comments, and I thank Stacy Schulman for assistance in this post.
By: Mike Horrocks The realities of the new economy and the credit crisis are driving businesses and financial institutions to better integrate new data and analytical techniques into operational decision systems. Adjusting credit risk processes in the wake of new regulations, while also increasing profits and customer loyalty will require a new brand of decision management systems to accelerate more precise customer decisions. There is a Webinar scheduled for Thursday that will insightfully show you how blending business rules, data and analytics inside a continuous-loop decisioning process can empower your organization - to control marketing, acquisition and account management activities to minimize risk exposure, while ensuring portfolio growth. Topics include: What the process is and the key building blocks for operating one over time Why the process can improve customer decisions How analytical techniques can be embedded in the change control process (including data-driven strategy design or optimization) If interested check out more - there is still time to register for the Webinar. And if you just want to see a great video - check out this intro.