Oscilar Team

Alternative Credit Scoring: Data, Models, and Governance

Posted

Posted

Oscilar Team
Contents

Share this article

Last updated: September 2026

TL;DR

Alternative credit scoring assesses creditworthiness using information outside the traditional bureau file. The argument about whether that information belongs in a credit decision is largely over, and the pages still making it are answering a question lenders stopped asking. What decides whether a programme works is duller and almost entirely unwritten: how each source is permissioned, what that obliges you to do about consent and coverage, what you have to re-test when a source changes, and whether the reasons your model produces can be given to a declined applicant in words they can act on.

What alternative credit scoring is

Alternative credit scoring is the assessment of creditworthiness using information that sits outside the traditional credit bureau file. The GAO describes that input as "information not used in traditional credit reporting" which "can be financial or nonfinancial in nature."

The distinction this article runs on is narrower than it first appears. Alternative scoring changes what the inputs are. It does not, by itself, change the policy engine, the testing regime, or the obligations attached to a lending decision. Programmes that treat it as a data project — source the feed, add the attributes, ship — meet the governance consequences later and out of order.

One scope note before going further. This is about lender-side practice: what a credit team has to build, test and defend. It is not about how a consumer improves their own file, which is a different subject with different advice.

Who it is for

Approximately 45 million Americans lack credit scores, as of the GAO's 2021 reporting. That population tends to be lower-income, younger, and disproportionately drawn from minority groups.

The figure is the most-quoted number in this field and it needs its date attached every time, which is why it carries one here. It describes a population as measured nearly five years ago, and population statistics move.

There is a boundary inside that number worth holding onto: thin file credit is not the same as no file at all. An applicant with two years of history and one tradeline can be scored badly by a conventional model. An applicant with no file at all cannot be scored by it. The first needs better inputs, the second needs different ones, and a programme that conflates them will underperform on both.

Where alternative data is actually used is worth being honest about, because the promotional literature implies it is everywhere. From fiscal years 2016 to 2020, the GAO found, only about 0.31% of FHA-insured loans went to borrowers without credit scores. Alternative data shows up far more in cards and smaller consumer loans than in mortgage, and the reasons are structural — secondary-market requirements and investor appetite rather than any reluctance to lend.

What actually counts as alternative data

Every page on this subject lists data types. The list is not the useful part, because two sources in the same category can carry entirely different obligations.

The GAO's own grouping is the starting point: bank account transactions; rental, utility and telecommunications payment history; and educational background and degree completion. Re-cut those by how each is permissioned and the picture becomes operationally usable, because permissioning is what determines your consent obligations, your coverage, and what breaks when a source changes.

Source type

What it evidences

Consent it requires

What happens when it changes

Consumer-permissioned — bank transaction data via an account connection

Income rhythm, expense structure, balance behaviour

Explicit, per-application, and revocable by the consumer

The connection drops silently. Coverage falls off first for the applicants least able to reconnect

Furnished — rental, utility and telecom payment reporting

A recurring obligation met, or not met, over time

Given to the furnisher, not to you; you inherit its scope

The furnisher changes reporting scope or cadence and your attribute quietly changes meaning

Inferred or third-party — derived scores and segments

Varies, and often a judgement rather than a fact

Frequently none you can point to on request

You may not be told at all

Read the right-hand column as the reason this cut matters. Consumer-permissioned data fails visibly and at the worst moment. Furnished data fails invisibly and changes what your model is measuring without changing your model. Inferred data can fail without notice, and it is the hardest to explain to an applicant or a regulator.

Educational attainment sits in the GAO's list and is treated separately later in this article, because it is the clearest case of a source whose problem is not permissioning at all.

How alternative signals enter a score

There are three routes, and the choice matters more for governance than for performance.

As attributes inside an existing scorecard. The new inputs join the existing model. This is the most invasive option: it changes the model everyone already relies on, and it means redeveloping and revalidating that model rather than adding to it.

As a separate model whose output feeds policy. A dedicated model consumes the alternative data and produces a score or a flag that policy then acts on. Alternative credit scoring models built this way are cleaner to govern, because the new component can be tested and monitored on its own.

As a second-look layer after a primary decline. The conventional decision runs first; declined applicants are reassessed using the alternative data. Second-look underwriting is the easiest of the three to govern and the best starting point, because the counterfactual is visible — you know exactly what the conventional process would have done, so the incremental effect of the new data is measurable rather than inferred.

What changes about testing in all three cases is the number of independent failure modes. Widening the input set does not simply add signal. Each new source has its own coverage profile, its own refresh cadence, its own definitional drift, and its own ways of going wrong — and those failures are not correlated with each other or with the bureau data you already understand.

Stability and what to re-test

Here is the change nobody schedules. A data source alters its coverage, its definitions, or its refresh cadence, and your model is now scoring something slightly different from what it was built on.

Nothing broke. No alert fired. The attribute simply means something else than it did last quarter.

The right response is to re-test the attribute, not merely to monitor the score. Score-level monitoring catches a large shift eventually, and by then the decisions have been made. Attribute-level testing catches the shift where it happened.

Two kinds of movement need separating, because they call for different actions. Population drift is your applicant mix changing — the model is fine, the inputs it is seeing are different. Source drift is the data itself changing under a stable population — the applicants are the same, the measurement moved. Treating the second as the first leads teams to recalibrate a model that did not need recalibrating.

And measure by segment. Aggregate performance can hold steady while a specific segment degrades badly, and the aggregate will read as stability. This is true of any model and it bites harder here, because the whole point of alternative data is to serve populations that are small in your existing book and therefore invisible in an aggregate.

Fair lending

This is the section that matters most, and it is guidance for your programme rather than a problem anyone can solve on your behalf.

The GAO gives a worked example: using educational attainment as an underwriting criterion could disadvantage populations with historically lower graduation rates. The variable may well be predictive. That is exactly why it is dangerous.

The general principle the example illustrates is the one to carry: a variable can be genuinely predictive and still function as a proxy variable for a protected characteristic, and predictiveness is not a defence. "The model found it useful" describes the mechanism of the problem, not a justification for it. Alternative data widens this exposure because it draws on behavioural and circumstantial information that correlates with protected characteristics in ways bureau data does not.

Testing for disparate impact means comparing outcomes across protected classes at the decision level, not the model-score level — approval rates, pricing, and terms as applicants actually experience them. It means doing that for each new attribute and each material model change, not once at launch. And it means having a documented process for what happens when a test finds something: whether the attribute is dropped, constrained, or retained with a justification someone is prepared to defend.

That work sits with your compliance and fair-lending functions, and it needs to be scoped before a source goes live rather than after. If you are evaluating any part of a credit stack and a supplier's answer to disparate-impact testing is that their platform handles it, ask precisely what is tested, against which classes, on whose data, and what the output looks like. The specific answers matter and they are frequently thinner than the claim.

Explainability and adverse action

A reason code has to be something an applicant can act on. That is the test, and it is stricter than producing a technically accurate explanation.

When a model reads deposit behaviour or utility payment history, its reasons have to be expressed in those terms and at that level. "Insufficient credit history" is a usable reason. "Your transaction pattern scored below threshold" is not — it is accurate and it tells the recipient nothing they can do anything about. The practical test: could you put this reason in a letter, and would the person receiving it know what to change?

Attribute-level explainability is therefore a selection criterion for a data source, not a downstream feature to add later. A source that yields only a composite score, with no ability to attribute a decision to a specific behaviour, constrains what you can lawfully tell a declined applicant.

The broader argument about AI-driven credit models and the scrutiny they attract is made in more depth in AI in credit underwriting, which covers reason-code generation and the regulatory frame around it. This article does not re-argue it.

Consent, permissioning and provenance

The GAO's privacy concern is direct: lenders gathering alternative data "without the knowledge or permission of consumers." Whatever a data supplier's terms assert about the permissions they hold, the reputational and regulatory exposure of a decision made on data an applicant did not know you had lands with the lender.

There is also a structural limitation in the GAO's analysis that constrains the whole opportunity. Consumers typically must opt in to credit products for such data to be included at all — which caps how much of the credit-invisible population any programme can reach. The applicants hardest to score are frequently the ones least likely to complete an optional connection step, so the reach of a programme is systematically smaller than the size of the population it targets.

Data provenance is the operational requirement that follows. For any attribute in a live decision, you should be able to state where it came from, under what permission, when it was obtained, and how long that permission runs. Programmes discover they cannot answer this at the point where someone asks — an examiner, a plaintiff's counsel, or an applicant exercising a data right.

Secondary use is the last piece and the easiest to get wrong. Data permissioned for one purpose is not automatically available for another. Bank transaction data a consumer connected for an income check is not thereby available for marketing segmentation, collections strategy, or a different product's underwriting. Each use needs its own basis.

How to evaluate a move to alternative data

Take the questions in this order, because each one constrains the next.

  1. Which population are you trying to reach? Thin-file existing applicants, no-file new applicants, or a segment you currently decline. These are different problems.

  2. Which sources actually cover that population? Coverage claims are usually quoted for the general population, not for the group you care about. Ask for coverage on your declined book.

  3. How is each source permissioned? This determines consent obligations, drop-off exposure, and what breaks when the source changes.

  4. What will you re-test, and when? Name the triggers before go-live: source change, coverage shift, segment degradation, model change.

  5. How will a decline be explained? If you cannot answer this at attribute level, the source is not usable regardless of its lift.

The scoping error to avoid is treating a change of inputs as though it were a change of decisioning platform, or the reverse. A rules engine applies policy to whatever inputs it is given; alternative scoring changes what the inputs are. They are complements, and confusing them is why these migrations get scoped wrong — teams replace a platform when they needed a data source, or bolt a data source onto a platform that cannot govern it.

A defensible first phase is usually second-look on a declined population. The volume is bounded, the counterfactual is measurable, the fair-lending analysis is tractable because you can compare against what the conventional process did, and a failure costs you an experiment rather than your primary decisioning path.

For the consumer lending case, B2C credit underwriting covers how Oscilar approaches the decisioning layer these inputs feed; B2B credit underwriting covers the commercial equivalent, where entity structure adds its own complications. If the question is the platform rather than the data, credit decisioning platform sets out what to evaluate. The companion treatment of bank transaction data specifically — cash flow underwriting, meaning what a lender actually reads in deposit behaviour and what has to happen to it before a policy can use it — is a natural next step from here.

Frequently asked questions

What is alternative credit scoring?

Alternative credit scoring assesses creditworthiness using information outside the traditional credit bureau file — bank account transactions, rental and utility payment histories, and similar sources. The GAO describes the input as information not used in traditional credit reporting, which may be financial or nonfinancial. It changes what a lender's decision is based on, not how the decision is governed.

What counts as alternative credit data?

The GAO groups it as bank account transactions; rental, utility and telecommunications payment histories; and educational background and degree completion. A more useful cut is by how each source is permissioned — consumer-permissioned, furnished by a provider, or inferred and third-party — because permissioning determines your consent obligations, your coverage, and what breaks when the source changes.

How does alternative credit scoring differ from credit scoring or a rules engine?

A rules engine applies policy to whatever inputs it is given. Alternative scoring changes what those inputs are. They are complements rather than substitutes, and confusing the two is why these projects get scoped wrong — a team replaces its decisioning platform when it needed a new data source, or adds a source to a platform that cannot govern it.

Who does alternative credit data actually help?

Primarily applicants a conventional model scores poorly or cannot score at all — roughly 45 million Americans lacked credit scores as of the GAO's 2021 reporting, a group tending to be lower-income, younger, and disproportionately from minority groups. The important distinction inside that is between thin-file applicants, who need better inputs, and no-file applicants, who need different ones.

How does alternative credit scoring work?

Alternative signals enter a decision by one of three routes: as attributes inside an existing scorecard, as a separate model whose output feeds policy, or as a second-look layer applied after a conventional decline. The second-look pattern is the easiest to govern and the usual starting point, because the counterfactual is visible and the incremental effect is measurable rather than inferred.

Does alternative data create fair-lending risk?

It can. The GAO's example is educational attainment, where use as an underwriting criterion could disadvantage populations with historically lower graduation rates. The principle generalises: a variable can be predictive and still act as a proxy for a protected characteristic, and predictiveness is not a defence. Testing for disparate impact has to happen at decision level, per attribute, and repeatedly.

What should lenders evaluate for explainability, testing and policy control?

For explainability, whether a decline can be attributed to a specific behaviour in terms an applicant can act on — not just a composite score. For testing, what gets re-tested when a source changes its coverage, definitions or cadence, and whether performance is measured by segment rather than in aggregate. For policy control, who can change a threshold, what evidence is required first, and whether the change is recorded.

What consent obligations come with alternative data?

They follow the permissioning model. Consumer-permissioned data requires explicit, revocable consent you hold directly. Furnished data carries consent given to the furnisher, whose scope you inherit rather than control. In all cases you should be able to state where an attribute came from, under what permission, and for how long — and data permissioned for one purpose is not automatically available for another.

Oscilar Team

The Oscilar Team is comprised of experts from many domains of risk operations. These articles express viewpoints and knowledge from a variety of sources and contributors across the organization.