What the system is actually using
People asked what a scoring system knows about them usually give one of two answers. Either it knows almost nothing, or it knows everything. Both are guesses, and the second one is a good deal more common since the phrase artificial intelligence entered general use.
The accurate answer is narrower and more troubling than either popular picture. These systems use whatever was recorded, and the record carries whoever held the pen: what got measured, about whom, and what never made it into the data at all.
Consider credit. A scoring model does not observe whether you repay debts. It observes a file: which of your obligations were reported, by whom, over what period, matched to your identity by whatever identifiers were available. Somebody with a long recorded history generates a file dense enough to support a confident prediction. Somebody who has borrowed informally, been paid in cash, or moved between addresses generates a thin one.
This is measurable and it has been measured. Across the records of 50 million consumers, model fit for predicting default was 16 per cent lower for the lowest income quartile than for the highest, and 18 per cent lower for minority borrowers. Approval rates for those groups sat around 35 per cent against over 50 per cent for others. The authors estimate that equalising the precision of scores would close roughly half the approval gap (Blattner and Nelson, 2021).
Read that carefully, because it is not the finding people expect. The claim is not primarily that the model is prejudiced. It is that the model is less certain about some people, and that uncertainty is not neutral in its effects. A lender facing a noisier prediction declines more often. The disadvantage is produced by having less recorded about you.
The same shape appears in lending prices. Across 5.7 million United States mortgages from 2009 to 2015, minority borrowers paid 4.7 to 4.9 basis points more on purchase loans for equivalent risk, costing over $450 million a year in aggregate. Algorithmic lenders showed 27 to 37 per cent smaller disparities than face-to-face lenders on one loan category in that period. By 2018 and 2019 there was no significant difference between them, and neither showed an advantage on the accept-or-reject decision at all (Bartlett, Morse, Stanton and Wallace).
"Algorithms discriminate less than humans" is a real finding from this paper, and it applies to one loan category, in one country, in the period 2009 to 2015, on price rather than on approval. The same authors found the advantage gone by 2018.
If you carry the first sentence out of this lesson and not the second, you have a fact that was true for six years and stopped being true, which is the most durable kind of wrong.
The practical version of this lesson is a question you can ask about any system that has scored you: what could it possibly have observed?
Most people, asked what a lending application on their phone knows about them, dramatically understate it. Kenya's data protection regulator, writing guidance for digital credit providers, described the position plainly: based on borrowers' acceptance of terms and conditions, these lenders' processes result in the collection of a large amount of customer information, including their call and message logs, phone information, and even their photographs and contacts. The regulator's guidance then prohibits accessing borrowers' phone records to gather third parties' contact details for debt collection, which tells you what had been happening (Office of the Data Protection Commissioner, Kenya, December 2023).
That regulator also fined a digital credit provider 2,975,000 shillings for using contact information obtained from third parties, without those people's consent, to send threatening messages during debt recovery. The provider challenged the penalty and a court dismissed the challenge.
Nigeria has a comparable registration regime for digital money lenders, created in 2022 after a joint regulatory investigation into possible violation of privacy and other rights. The framework exists and is verifiable. The numbers attached to it in circulation, counts of licensed, approved and delisted applications, could not be verified against a primary document, so this course does not repeat them.
Your cousin is declined by a lending application after twenty seconds. He has never missed a payment on anything and concludes the system is prejudiced against him. What is the most likely explanation on the evidence?
What does the Kenyan regulator's guidance to digital lenders establish?
A colleague says the research shows algorithmic lenders discriminate less than human ones, so automating a decision makes it fairer. What is the accurate correction?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.