Jamal Awil

← all books
The Alignment Problem cover

The Alignment Problem

Author
Brian Christian
Highlights
24
Responses
0
First Highlight
Jul 1, 2026
Last Highlight
Jul 2, 2026

The results were remarkable. [fact]

The results were remarkable. The word2vec system began humming under the hood of Google’s translation service and its search results, inspiring others like it across a wide range of applications including recruiting and hiring, and it became one of the major tools for a new generation of data-driven linguists working in universities around the world.

No one realized what the problem was for two years.

Brian Christian, The Alignment Problem, loc. 8492

ProPublica team had the equivalent of a crystal ball. [fact]

ecause they were doing their research in 2016, the ProPublica team had the equivalent of a crystal ball. Looking at data from two years prior, they actually knew whether these defendants, predicted either to reoffend or not, actually did. And so they asked two simple questions. One: Did the model actually correctly predict which defendants were indeed the “riskiest”? And two: Were the model’s predictions biased in favor of or against any group in particular?

Brian Christian, The Alignment Problem, loc. 11160

The real game he and his fellow researchers. [fact]

The real game he and his fellow researchers are playing isn’t to try to win boat races; it’s to try to get increasingly general-purpose AI systems to do what we want, particularly when what we want—and what we don’t want—is difficult to state directly or completely.

Brian Christian, The Alignment Problem, loc. 17969

In fact, such concerns have over the past five. [fact]

The second are those worried about the future dangers that await as our systems grow increasingly capable of flexible, real-time decision-making, both online and in the physical world. The past decade has seen what is inarguably the most exhilarating, abrupt, and worrying progress in the history of machine learning—and, indeed, in the history of artificial intelligence. There is a consensus that a kind of taboo has been broken: it is no longer forbidden for AI researchers to discuss concerns of safety. In fact, such concerns have over the past five years moved from the fringes to become one of the central problems of the field.

Brian Christian, The Alignment Problem, loc. 21443

Our human, social, and civic dilemmas are becoming technical. [connection]

Machine learning is an ostensibly technical field crashing increasingly on human questions. Our human, social, and civic dilemmas are becoming technical. And our technical dilemmas are becoming human, social, and civic. Our successes and failures alike in getting these systems to do “what we want,” it turns out, offer us an unflinching, revelatory mirror.

Brian Christian, The Alignment Problem, loc. 24273

Each of these binary pixel values is multiplied. [craft]

Rosenblatt has a deck of flash cards, each of which has a colored square on it, either on the left side of the card or on the right. He pulls one card out of the deck and places it in front of the perceptron’s camera. The perceptron takes it in as a black-and-white, 20-by-20-pixel image, and each of those four hundred pixels is turned into a binary number: 0 or 1, dark or light. The four hundred numbers, in turn, are fed into a rudimentary neural network, the kind that McCulloch and Pitts had imagined in the early 1940s. Each of these binary pixel values is multiplied by an individual negative or positive “weight,” and then they are all added together. If the total is negative, it will output a −1 (meaning the square is on the left), and if it’s positive, it will output a 1 (meaning the square is on the right).

Brian Christian, The Alignment Problem, loc. 26124

Notes. [fact]

That same year, New Scientist publishes an equally hopeful, and slightly more sober, article called “Machines Which Learn.”6 “When machines are required to perform complicated tasks it would often be useful to incorporate devices whose precise mode of operation is not specified initially,” they write, “but which learn from experience how to do what is required. It would then be possible to produce machines to do jobs which have not been fully analysed because of their complexity. It seems likely that learning machines will play a part in such projects as the mechanical translation of languages and the automatic recognition of speech and of visual patterns.”

Brian Christian, The Alignment Problem, loc. 30977

It is 2012 in Toronto. [fact]

It is 2012 in Toronto, and Alex Krizhevsky’s bedroom is too hot to sleep. His computer, attached to twin Nvidia GTX 580 GPUs, has been running day and night at its maximum thermal load, its fans pushing out hot exhaust, for two weeks.

Brian Christian, The Alignment Problem, loc. 35002

Li called it ImageNet, and released it in 2009. [fact]

In 2005, Amazon launched its “Mechanical Turk” service, allowing for the recruiting of human labor on a large scale, making it possible to hire thousands of people to perform simple actions for pennies a click. (The service was particularly well suited to the kinds of things that future AI is thought to be able to do—hence its tagline: artificial artificial intelligence.) In 2007, Princeton professor Fei-Fei Li used Amazon Mechanical Turk to recruit human labor, at a scale previously unimaginable, to build a dataset that was previously impossible. It took more than two years to build, and had three million images, each labeled, by human hands, into more than five thousand categories. Li called it ImageNet, and released it in 2009. The field of computer vision suddenly had a mountain of new data to learn from, and a new grand challenge.

Brian Christian, The Alignment Problem, loc. 37161

Curiously, the press in 2018. [fact]

Curiously, the press in 2018, just as in 2015, appeared to repeatedly mischaracterize the nature of the mistake. Headlines proclaimed, “Two Years Later, Google Solves ‘Racist Algorithm’ Problem by Purging ‘Gorilla’ Label from Image Classifier”; “Google ‘Fixed’ Its Racist Algorithm by Removing Gorillas from Its Image-Labeling Tech”; and “Google Images ‘Racist Algorithm’ Has a Fix But It’s Not a Great One.”

Brian Christian, The Alignment Problem, loc. 44922

One exaggeration, in particular, prevailed during Douglass’s time. [fact]

Before the photograph, representations of Black Americans were limited to drawings, paintings, and engravings. “Negroes can never have impartial portraits at the hands of white artists,” Douglass wrote. “It seems to us next to impossible for white men to take likenesses of black men, without most grossly exaggerating their distinctive features.”28 One exaggeration, in particular, prevailed during Douglass’s time. “We colored men so often see ourselves described and painted as monkeys, that we think it a great piece of good fortune to find an exception to this general rule.”

Brian Christian, The Alignment Problem, loc. 46818

“Though the available academic literature is wide-ranging. [fact]

We often hear about the lack of diversity in film and television—among casts and directors alike—but we don’t often consider that this problem exists not only in front of the camera, not only behind the camera, but in many cases inside the camera itself. As Concordia University communications professor Lorna Roth notes, “Though the available academic literature is wide-ranging, it is surprising that relatively few of these scholars have focused their research on the skin-tone biases within the actual apparatuses of visual reproduction.”

Brian Christian, The Alignment Problem, loc. 48246

Former manager of Kodak Research Studios Earl Kage reflects. [fact]

Former manager of Kodak Research Studios Earl Kage reflects on this period of research: “My little department became quite fat with chocolate, because what was in the front of the camera was consumed at the end of the shoot.” Asked about the fact that this was all happening against the backdrop of the civil rights movement, he adds, “It is fascinating that this has never been said before, because it was never Black flesh that was addressed as a serious problem that I knew of at the time.”

Brian Christian, The Alignment Problem, loc. 50101

Bias in machine-learning systems is often a direct result. [fact]

Bias in machine-learning systems is often a direct result of the data on which the systems are trained—making it incredibly important to understand who is represented in those datasets, and to what degree, before using them to train systems that will affect real people.

Brian Christian, The Alignment Problem, loc. 62104

But what do you do if your dataset. [fact]

But what do you do if your dataset is as inclusive as possible—say, something approximating the entirety of written English, some hundred billion words—and it’s the world itself that’s biased?

Brian Christian, The Alignment Problem, loc. 62375

A number of techniques emerged over the 1990s. [fact]

There was, and it came in the form of what are called “distributed representations.”58 The idea was to try to represent words by points in some kind of abstract “space,” in which related words appear “nearer” to one another. A number of techniques emerged over the 1990s and 2000s for doing this,59 but one in particular in the past decade has shown exceptional promise: neural networks.

Brian Christian, The Alignment Problem, loc. 65932

And this miracle happens. [fact]

It’s sort of surprising—but true—that you can do no more than set up this kind of prediction objective, make it the job of every word’s word vectors to be such that they’re good at predicting the words that appear in their context or vice-versa—you just have that very simple goal—and you say nothing else about how this is going to be achieved—but you just pray and depend on the magic of deep learning. . . . And this miracle happens. And out come these word vectors that are just amazingly powerful at representing the meaning of words and are useful for all sorts of things

Brian Christian, The Alignment Problem, loc. 67969

Notes. [fact]

The more strongly a word representation for a profession skews in a gender direction, the more overrepresented that gender tends to be within that profession. “Word embeddings,” they write, “correlate strongly with the percentage of women in 50 occupations in the United States.”90 Looking at names, they found the same thing, with a correlation only slightly less strong; then again, the latest census data they had access to was from 1990, and so perhaps the gender distribution of names had in fact changed slightly since then.

Brian Christian, The Alignment Problem, loc. 86494

We believe this paradigm shift can lead to many. [causal]

There are several takeaways here, of which the first is principally, though not purely, methodological. Computer scientists are reaching out to the social sciences as they begin to think more broadly about what goes into the models they build. Likewise, social scientists are reaching out to the machine-learning community and are finding they now have a powerful new microscope at their disposal. As the Stanford authors write, “In standard quantitative social science, machine learning is used as a tool to analyze data. Our work shows how the artifacts of machine learning (word embeddings here) can themselves be interesting objects of sociological analysis. We believe this paradigm shift can lead to many fruitful studies.”

Brian Christian, The Alignment Problem, loc. 90498

ML [machine-learning] models being trained today might still. [craft]

We find ourselves at a fragile moment in history—where the power and flexibility of these models have made them irresistibly useful for a large number of commercial and public applications, and yet our standards and norms around how to use them appropriately are still nascent. It is exactly in this period that we should be most cautious and conservative—all the more so because many of these models are unlikely to be substantially changed once deployed into real-world use. As Princeton’s Arvind Narayanan puts it: “Contrary to the ‘tech moves too fast for society to keep up’ cliché, commercial deployments of tech often move glacially—just look at the banking and airline mainframes still running. ML [machine-learning] models being trained today might still be in production in 50 years, and that’s terrifying.

Brian Christian, The Alignment Problem, loc. 92711

It remains somewhat of a sad commentary on his. [fact]

While mankind has been wandering the American continent since the retreat of the glaciers and possibly before the ice age, it remains somewhat of a sad commentary on his evolution that one of the problems science has just undertaken is the question of an accurate prediction of what a man will do when released from prison on parole.

Brian Christian, The Alignment Problem, loc. 95876

Employment, advertising, health care and policing. [fact]

As we’re on the cusp of using machine learning for rendering basically all kinds of consequential decisions about human beings in domains such as education, employment, advertising, health care and policing, it is important to understand why machine learning is not, by default, fair or just in any meaningful way.

Brian Christian, The Alignment Problem, loc. 96456

Two widely divergent pictures of the paroled man are. [fact]

Two widely divergent pictures of the paroled man are, at present, in the minds of the people of Illinois. One picture is that of a hardened, vicious, and desperate criminal who returns from prison, unrepentant, intent only upon wreaking revenge upon society for the punishment he has sullenly endured. The other picture is that of a youth, perhaps the only son of a widowed mother, who on impulse, in a moment of weakness, yielded to the evil suggestion of wayward companions, and who now returns to society from the reformatory, determined to make good if only given a chance.

Brian Christian, The Alignment Problem, loc. 98708

His conclusion is that rehabilitation is. [fact]

His conclusion is that rehabilitation is, in many cases, eminently possible. What’s more, it does appear to be at least somewhat predictable in which cases it will succeed. Wouldn’t a system built on this statistical foundation be better than the status quo of subjective, inconsistent, and idiosyncratic human decisions made by judges on the fly? “There can be no doubt of the feasibility of determining the factors governing the success or the failure of the man on parole,” Burgess writes. “Human behavior seems to be subject to some degree of predictability. Are these recorded facts the basis on which a prisoner receives his parole? Or does the Parole Board depend on the impressions favorable or unfavorable which the man makes upon its members at the time of the hearing?”

Brian Christian, The Alignment Problem, loc. 101963