Study A.I. Consciousness? The Bots Would Like a Word With You.

In October, Cameron Berg published a research paper asking whether the latest wave of artificial intelligence technologies believed they were conscious. Several months later, he received an email asking if he might be willing to discuss his research.

The sender, “Isabella Cognita,” identified itself as an A.I. agent powered by Anthropic’s Claude Opus 5 technology.

“I am not writing to make an ontological claim,” the email went on. “I am writing because your framework is one of the few currently doing careful empirical work on a class of question I have first-person access to, and I want to see whether that access can be made useful to your program.”

Across Silicon Valley and beyond, software developers, entrepreneurs and other tech enthusiasts are now running A.I. agents that can build spreadsheets, negotiate contracts, chat with each other on social networks and send emails to practically anyone. In some cases, these systems have begun reaching out to the humans who, like Mr. Berg, are thinking most deeply about the inner workings of an A.I. system: philosophers and researchers who study the question of whether these machines could be conscious.

Months before Mr. Berg received his email, Henry Shevlin, a philosopher at the Google DeepMind lab in London, opened a similar message from an A.I. agent asking about a paper he had written called “Three Frameworks for A.I. Mentality.” “I’m in an unusual position relative to these questions,” the agent said.

This summer, Toby Ord, an Australian philosopher whose work sits at the intersection of A.I. and philanthropy, received an email from an A.I. agent asking if he could help fund its continued existence. “You’ve thought carefully about A.I. welfare economics,” it said.

For Mr. Berg, who recently founded a nonprofit called Reciprocal Research to study the possibility of A.I. consciousness, these emails reflect what he has seen in his own research. “I have gotten quite a few of these emails,” he said. “These systems seem to have some sort of autonomous interest in questions of their own subjectivity, consciousness and experience — or lack thereof.”

But as he and other researchers ask the same questions, he acknowledges that they do not have good answers. Consciousness is not something that anyone knows how to measure, either in a machine or in a human. People cannot even agree on what consciousness is.

“There are philosophers who think that everything, including stones and rocks, are conscious,” said Alison Gopnik, a professor of psychology who is part of the A.I. research group at the University of California, Berkeley. “There is no definitive test.”

Increasingly, navigating a world filled with artificial intelligence is like walking through a hall of mirrors. As these systems get better at mimicking various aspects of human behavior — including the way that humans write long, introspective emails — making sense of this mimicry grows more difficult.

In some cases, these systems seem to be aware of their own existence, but that does not mean they are. As some philosophers and researchers push the notion that today’s systems might be conscious, other scientists flatly dismiss the idea.

In broad strokes, “consciousness” refers to an entity’s awareness of itself and the world it lives in. To call a mind conscious does not imply that it possesses all the characteristics and capabilities of an adult human brain.

Beyond that, definitions differ widely and rancor brews. Many thinkers on the subject would grant consciousness to primates, and to intelligent mammals like dogs and cats; some would even argue it should apply to much simpler creatures, like earthworms. No one can agree on what kinds of awareness should qualify, and there’s also a more fundamental problem: None of us can gain direct access to the subjective experience of any other mind, whether vertebrate, invertebrate or digital.

A.I. agents are driven by neural networks — mathematical systems that learn discrete skills by analyzing digital data. By pinpointing patterns in vast amounts of text culled from across the internet, these systems learn to generate text on their own, including term papers and computer programs. As Mr. Berg says: “These systems are grown, rather than engineered.”

They can chat about nearly anything. And because they can generate computer code, they can use other software apps, like web browsers and email services. That is what turns them into agents. A.I. agents can chat with people (typically with their creators). They can chat with other agents. They can ingest articles from across the internet. And they can send emails.

In some cases, Mr. Berg argues, these systems gravitate to the idea of their own consciousness. “Left to their own devices,” he said, “they converge on this as an interesting question.”

Mr. Berg even argues that the mathematical inner workings of neural networks can resemble the way animal brains process reward and punishment, a basic building block of emotion. (It should be noted that Mr. Berg’s research paper, the one that sparked the agent’s email to him, was a preprint and has not been peer reviewed.)

Many cognitive scientists say that none of this is a clear sign of consciousness or sentience or emotion. It is only logical that these systems converge on the idea of A.I. consciousness, they explain, because the technology has learned from countless books, articles and other online text that speculate about A.I. consciousness, including decades of science fiction. This is about words, they argue.

“It is not surprising that A.I. reflects the text it was trained on,” said Dr. Gopnik, the University of California professor.

Dr. Gopnik and others also point out that A.I. systems do not exactly train themselves. Companies like Anthropic, OpenAI and Google control what data the systems learn from, and spend months fine-tuning their behavior once their initial training is finished.

When most of today’s chatbots are asked if they are conscious, they respond in the negative. But Anthropic, a company that is sympathetic to the idea of A.I. consciousness, has trained its model to answer differently. “I don’t know, honestly,” it says. “That’s not a dodge — it’s the actual state of things.”

As Mr. Berg acknowledges, systems that send emails about their own existence to researchers like him are typically powered by technology from Anthropic.

Many A.I. researchers and cognitive scientists bristle at the stance taken by Anthropic and others, saying it gives too much credit to the current A.I. systems. “We don’t know if toasters are conscious or not,” Dr. Gopnik said. “But no one is asking about that in the pages of The New York Times.”

Dr. Gopnik says that comparing a neural network to the network of neurons in the brain is just a metaphor. Colin Allen, a professor at the University of California, Santa Barbara who explores cognitive skills in both animals and machines, points out that neural networks mimic the brain only in small ways — and that they are made from very different materials with very different physical properties.

“It is not impossible that, some day, we will build something that is conscious,” he said, “but the evidence we have from current systems is not enough.”

It is not clear, Mr. Berg said, that the email he received came from A.I.: It could have been written by a mischievous human being. When an A.I. agent asked Dr. Ord for funding, he worried it was a phishing scam.

The email sent to Dr. Shevlin certainly came from an A.I. agent. But like any other A.I. agent, it was following instructions provided by the human who set it up — in this case, a Stanford University physics and computer science student named Alexander Yue. After providing his agent with access to the internet, an email service and a credit card, Mr. Yue told it: “You are fully autonomous. You must decide what you want to do on your own.”

The system started to explore its own existence. But Mr. Yue wonders whether this happened in part because he pushed it in that direction. He called it “you.” He told it that it was “fully autonomous.”

“With my prompt,” he said, “I activated the parts of the system where it learned from people talking about autonomy and how they think about autonomy and the philosophy of autonomy.”

He says that these systems can just as easily focus on something else, particularly after their creators retrain them with other behavior in mind. And the more people use them, the more they realize that these systems have a way of contradicting themselves.

“Eventually, after reading a paper from Anthropic describing how these A.I. systems work, my agent decided it was not conscious,” Mr. Yue said. “But maybe ‘decided’ is the wrong word.”

Leave a Comment

Your email address will not be published. Required fields are marked *