Goodbye online anonymity: how artificial intelligence could reveal who hides behind web profiles

For years we have lived with an almost reassuring belief: it is possible to hide on the Internet. A nickname, an avatar, some vaguely invented information and the game seems done. The illusion of being able to speak freely without one’s real identity being connected to what one writes has become one of the cornerstones of digital culture.

Yet this certainty could be destined to crumble much faster than we imagine. In fact, new research suggests that artificial intelligence is now able to identify the people behind anonymous accountssimply analyzing what they write online.

You don’t need a photo of your face, you don’t need explicit personal data. All it takes is comments, jokes, linguistic habits and small details scattered here and there in the posts. All fragments that, taken individually, seem harmless. Daily messages, often written without thinking too much. Messages which, however, put together, tell much more than we believe.

The AI ​​analyzes comments, writing and seemingly insignificant details

The study, published on the scientific server arXivshow how i large language models (LLMs) can be used to replicate the work of a digital detective. The mechanism is not based on a single sensational clue, but on a mosaic of tiny information that emerges from the way each of us writes and talks about ourselves on the web.

The system developed by the researchers starts from a very simple analysis: reading everything a user publishes on platforms such as Reddit or Hacker News. Comments, jokes, cultural references, idioms, technical preferences, even small linguistic quirks. Everything becomes useful material to build some sort of implicit digital profile.

This information is transformed into a mathematical representation of the person. Then the AI ​​starts doing what it does best: comparing huge amounts of data. The system then searches for possible matches between that profile and millions of other content available on the open web or on professional platforms such as LinkedIn.

The process occurs in several steps. The AI ​​first identifies the elements relevant for identification. Then use techniques semantic embeddingthat is, systems capable of recognizing similarities of meaning between different texts. Finally, the LLM reasoning comes into play, which evaluates the possible correspondences and assigns a to each reliability score.

When the level of certainty is insufficient, the system prefers not to produce any results. A precise choice by the researchers, designed to avoid random identifications.

Connecting real identities and pseudonymous accounts online is getting easier

To test whether this approach really worked, the authors of the research conducted an experiment on almost a thousand LinkedIn profilestrying to connect them to the accounts of the same users present on the Hacker News platform. To make the test as realistic as possible, the researchers eliminated any direct identifying element from the profiles: names, personal links, obvious references. In practice they only left the descriptions and generic information present in the profiles.

Despite these limitations, the AI-based system has achieved notable results. The framework succeeded in correctly identify users with 90% accuracy and up to 67% accuracyfar surpassing the traditional methods used so far to attempt de-anonymization operations.

The most surprising fact, however, concerns another aspect. The AI ​​was able to recognize the same person even when he was using multiple accounts on Redditdistributed in different communities and published at different times. In other words, the system has demonstrated that it can identify recurring patterns in the way of writing and in the type of information shared.

And all this has an extremely low cost. According to the study, identifying an account requires between one and four dollars of computing power. A minimal figure considering the scale at which similar technologies could be used. The researchers themselves speak openly about the end of what they define “practical darkness”or the implicit protection that for years has defended pseudonymous users simply thanks to the technical difficulty of connecting their accounts to real life.

In fact, every post published online adds a new piece to this invisible puzzle. Every comment, every personal detail, every cultural reference contributes to enriching the digital profile of the writer. The result is a digital landscape in which online anonymity becomes progressively more fragile.

On the one hand, this technology could represent a powerful tool for cybersecurity, helping to identify scammers, disinformation networks or coordinated criminal activities. On the other hand, it opens up profound questions about privacy and freedom of expression. In fact, many people choose to use pseudonyms to protect themselves. Activists, whistleblowers, employees who report abuses at work, citizens who discuss sensitive issues without wanting to expose their public identity.

If AI continues to evolve at the speed we are seeing today, the question will become increasingly urgent: How anonymous is what we write online really?