14.1 Predictive Analytics
1.1 The line that matters
| Predicting things | Predicting people |
|---|---|
| the weather | whether a convict is likely to reoffend |
| the spread of diseases | whether a loan applicant will default |
| whether an insurance customer will make expensive claims |
“The latter have a direct effect on individual people’s lives.”
Why organizations are structurally biased toward "no":
"Payment networks want to prevent fraudulent transactions, banks want to avoid bad loans, airlines want to avoid hijackings, companies want to avoid hiring ineffective people. FROM THEIR POINT OF VIEW, THE COST OF A MISSED BUSINESS OPPORTUNITY IS LOW, BUT THE COST OF A BAD LOAN OR A PROBLEMATIC EMPLOYEE IS MUCH HIGHER, so it is expected for organizations to want to be cautious. IF IN DOUBT, THEY ARE BETTER OFF SAYING NO."
And why that becomes catastrophic in aggregate:
"As algorithmic decision making becomes more widespread, someone who has (ACCURATELY OR FALSELY) BEEN LABELED AS RISKY may suffer A LARGE NUMBER OF THOSE 'NO' DECISIONS. Systematically being excluded from JOBS, AIR TRAVEL, INSURANCE COVERAGE, PROPERTY RENTAL, FINANCIAL SERVICES, and other key aspects of society is such a large constraint of an individual's freedom that it has been called 'ALGORITHMIC PRISON.'"
"In countries that respect human rights, THE CRIMINAL JUSTICE SYSTEM PRESUMES INNOCENCE UNTIL PROVEN GUILTY; ON THE OTHER HAND, AUTOMATED SYSTEMS CAN SYSTEMATICALLY AND ARBITRARILY EXCLUDE A PERSON FROM PARTICIPATING IN SOCIETY WITHOUT ANY PROOF OF GUILT, AND WITH LITTLE CHANCE OF APPEAL."
1.2 Bias and discrimination
The hope, stated fairly first:
"Decisions made by an algorithm are NOT NECESSARILY ANY BETTER OR ANY WORSE than those made by a human. EVERY PERSON IS LIKELY TO HAVE BIASES, EVEN IF THEY ACTIVELY TRY TO COUNTERACT THEM, and discriminatory practices can become CULTURALLY INSTITUTIONALIZED. THERE IS HOPE that basing decisions on data could be MORE FAIR and give a better chance to people who are often overlooked or disadvantaged."
Then the mechanism by which it fails:
"When we develop predictive analytics and AI systems, WE ARE NOT MERELY AUTOMATING A HUMAN'S DECISION by using software to specify the rules; WE ARE LEAVING THE RULES THEMSELVES TO BE INFERRED FROM DATA. However, the patterns learned are OPAQUE: EVEN IF THE DATA INDICATES A CORRELATION, WE MAY NOT KNOW WHY."
"IF THE INPUT TO AN ALGORITHM CARRIES A SYSTEMATIC BIAS, THE SYSTEM WILL MOST LIKELY LEARN AND AMPLIFY THAT BIAS IN ITS OUTPUT."
The proxy problem — why "just don't use protected attributes" doesn't work:
Anti-discrimination laws prohibit treating people differently based on Protected traits: ethnicity, age, gender, sexuality, disability, beliefs.
But: “Other features may be analyzed — What happens if they are correlated with protected traits? For example, In racially segregated neighborhoods, a person’s postal code or even their IP address is A strong predictor of race.”
⇒ “Put like this, It seems ridiculous to believe that an algorithm could somehow take biased data as input and produce fair and impartial output from it. yet this belief often seems to be implied by proponents of data-driven decision making — an attitude that has been satirized as ‘Machine learning is like money laundering for bias.’”
"PREDICTIVE ANALYTICS SYSTEMS MERELY EXTRAPOLATE FROM THE PAST; IF THE PAST IS DISCRIMINATORY, THEY CODIFY AND AMPLIFY THAT DISCRIMINATION. IF WE WANT THE FUTURE TO BE BETTER THAN THE PAST, MORAL IMAGINATION IS REQUIRED, AND THAT'S SOMETHING ONLY HUMANS CAN PROVIDE. DATA AND MODELS SHOULD BE OUR TOOLS, NOT OUR MASTERS."
1.3 Responsibility and accountability
"If a human makes a mistake, THEY CAN BE HELD ACCOUNTABLE, and the person affected CAN APPEAL. Algorithms make mistakes too, BUT WHO IS ACCOUNTABLE IF THEY GO WRONG?"
- "When a self-driving car causes an accident, WHO IS RESPONSIBLE?"
- "If an automated credit scoring algorithm systematically discriminates against people of a particular race or religion, IS THERE ANY RECOURSE?"
- "If a decision by your ML system comes under judicial review, CAN YOU EXPLAIN TO THE JUDGE HOW THE ALGORITHM MADE ITS DECISION?"
"PEOPLE SHOULD NOT BE ABLE TO EVADE THEIR RESPONSIBILITY BY BLAMING AN ALGORITHM."
Credit scores vs predictive scoring — the crucial distinction:
| Traditional credit score | ML-based predictive scoring | |
|---|---|---|
| Question it answers | "HOW DID YOU BEHAVE IN THE PAST?" | "WHO IS SIMILAR TO YOU, AND HOW DID PEOPLE LIKE YOU BEHAVE IN THE PAST?" |
| Basis | Relevant facts about a person's ACTUAL BORROWING HISTORY | A MUCH WIDER RANGE OF INPUTS, MUCH MORE OPAQUE |
| Correction | Errors in the record CAN BE CORRECTED (although agencies "normally do not make this easy") | "If a decision is incorrect because of ERRONEOUS DATA, RECOURSE IS ALMOST IMPOSSIBLE" |
| Ethical problem | — | "Drawing parallels to others' behavior IMPLIES STEREOTYPING PEOPLE — for example, based on where they live (A CLOSE PROXY FOR RACE AND SOCIOECONOMIC CLASS). WHAT ABOUT PEOPLE WHO GET PUT IN THE WRONG BUCKET?" |
The statistical fallacy applied to individuals:
"Much data is STATISTICAL in nature, which means that EVEN IF THE PROBABILITY DISTRIBUTION ON THE WHOLE IS CORRECT, INDIVIDUAL CASES MAY WELL BE WRONG. If the average life expectancy in your country is 80 years, THAT DOESN'T MEAN YOU'RE EXPECTED TO DROP DEAD ON YOUR 80TH BIRTHDAY. From the average and the distribution, YOU CAN'T SAY MUCH ABOUT THE AGE TO WHICH ONE PARTICULAR PERSON WILL LIVE."
"A BLIND BELIEF IN THE SUPREMACY OF DATA FOR MAKING DECISIONS IS NOT ONLY DELUSIONAL BUT ALSO POSITIVELY DANGEROUS."
The dual-use problem, stated without flinching:
"Analytics can reveal financial and social characteristics of people's lives. ON THE ONE HAND, this power could be used to FOCUS AID AND SUPPORT TO HELP THOSE WHO NEED IT MOST. ON THE OTHER HAND, IT IS SOMETIMES USED BY PREDATORY BUSINESSES SEEKING TO IDENTIFY VULNERABLE PEOPLE AND SELL THEM RISKY PRODUCTS SUCH AS HIGH-COST LOANS OR WORTHLESS COLLEGE DEGREES."
1.4 Feedback loops
Echo chambers:
"When services become good at predicting the content users want to see, THEY MAY END UP SHOWING PEOPLE ONLY OPINIONS THEY ALREADY AGREE WITH, leading to ECHO CHAMBERS in which STEREOTYPES, MISINFORMATION, AND POLARIZATION CAN BREED. WE ARE ALREADY SEEING THE IMPACT SOCIAL MEDIA ECHO CHAMBERS CAN HAVE ON ELECTION CAMPAIGNS."
The self-reinforcing spiral — trace it:
Employers use credit scores to evaluate potential hires. Trace what that does to someone whose luck turns.
“It’s a downward spiral due to poisonous assumptions, hidden behind a camouflage of mathematical rigor and data.”
A second, non-obvious example:
"Economists found that when GAS STATIONS IN GERMANY introduced ALGORITHMIC PRICES, COMPETITION WAS REDUCED AND PRICES FOR CONSUMERS WENT UP BECAUSE THE ALGORITHMS LEARNED TO COLLUDE." (Nobody programmed collusion. It emerged.)
The tool for anticipating this — systems thinking:
"We can't always predict when such feedback loops may happen. However, MANY CONSEQUENCES CAN BE PREDICTED BY THINKING ABOUT THE ENTIRE SYSTEM — NOT JUST THE COMPUTERIZED PARTS, BUT ALSO THE PEOPLE INTERACTING WITH IT — an approach known as SYSTEMS THINKING."
"DOES THE SYSTEM REINFORCE AND AMPLIFY EXISTING DIFFERENCES BETWEEN PEOPLE (e.g., MAKING THE RICH RICHER OR THE POOR POORER), OR DOES IT TRY TO COMBAT INJUSTICE? EVEN WITH THE BEST INTENTIONS, WE MUST BEWARE OF THE POSSIBILITY OF UNINTENDED CONSEQUENCES."