1 From being seen to being understood: New changes in the personal digital existence in the AI era
In human society, being "recognized" has always been a scarce resource.
Before the internet, a person's fame often depended on whether they were part of an existing communication system: this could be traditional media coverage, the influence of their organization, or interpersonal communication in the real world. For most ordinary people, even with unique abilities and experience, it was difficult to break through the boundaries of communication in the real world.
The internet age once showed many people a new possibility: anyone can publish content online, allowing their thoughts, experiences, and works to enter a wider world, and in the long run, form their own digital existence.
While the internet lowered the barrier to expression, it didn't completely solve the problem of individuals being recognized. In the search engine era, the basic unit of information organization remained the webpage. The core task of search engines was to help users find relevant content from a large number of pages, not to determine who was behind those pages.
Therefore, an individual's digital visibility often depends on whether their content can continuously accumulate online influence through long-term dissemination, external citations, and community discussions. As this information continues to spread, people eventually recognize not just a few articles, but the authors behind them.
However, this path still presents a high barrier for most ordinary people: even if a person has rich experience and unique perspectives, if there is a lack of sufficient dissemination, this content can easily be scattered across different pages, platforms, and time points, making it difficult to break through the original network boundaries and further connect into a stable personal identity.
Besides relying on content dissemination to accumulate influence, there was another way to establish a personal identity in the traditional internet era:Enter a knowledge system with public credit attributes. For example, open knowledge platforms like Wikipedia create separate entries for individuals with public influence through community rules and editing mechanisms. This approach essentially allows a credible knowledge system to verify a person's identity, experience, and achievements.
However, this approach also has a high barrier to entry—biographical entries typically require sufficiently independent and reliable sources, as well as verifiable public influence. In other words, an individual often needs to gain recognition from the public communication system before they can more easily enter this centralized knowledge system. However, for the vast majority of ordinary people, even with rich experience and long-term accumulation, it is difficult to establish sufficient public influence through traditional methods.
However, the development of AI is changing this situation. During training and information acquisition, AI models encounter a large amount of publicly available text from the internet, including open-source communities, technical forums, blogs, documents, and other publicly available materials. For individuals, as long as they leave accessible and citationable information on the open internet over a long period, this content has the potential to become one of the information sources that AI references when understanding the world.
Compared to search engines, which rely more on page ranking, traffic, and external links, AI uses a more diverse approach to assessing the value of information. This allows even small-scale, but long-accumulated and uniquely valuable, personal content to be understood and connected.
On the other hand, the way people interact with internet information is changing in the AI era: In the search engine era, users needed to actively break down problems, find pages through keywords, and then read, filter, and organize information themselves; while in the AI era, users are increasingly accustomed to directly stating their goals and letting AI help them complete information retrieval, summarization, and analysis.
Therefore, users' questions have gradually shifted from searching for "what content is relevant to my problem" to judging "who is behind this content, and whether this person is trustworthy," such as: Who are the long-term practitioners in this field? Are there any trustworthy experts? What are this person's views and experiences?
To answer these questions, AI will need to process no more than just a single webpage, but...Information and its relationships scattered across the InternetOnly when a person's experience, works, viewpoints, professional field, and long-term behavior are considered...Can be attributed to the same personOnly then can AI gradually form a holistic understanding of this person.
Furthermore, if this information can further reflect a person's knowledge domain, way of thinking, and long-term practice path, then AI will no longer just understand "who this person is," but will be able to form a complete understanding of this person's characteristics.
This means that in the AI era, the way individuals are perceived by the digital world is changing: a person no longer needs to be a public figure in the traditional sense, have a large following, or enter a centralized knowledge system to have a chance to form a clear public identity. As long as one continuously leaves behind authentic experiences, unique perspectives, practical records, and a stable accumulation of knowledge on the open internet, these scattered digital traces can be reconnected and organized by AI, gradually forming a coherent digital personality about that person.
2. In the era of platform-based internet, why is it becoming increasingly difficult to form a personal digital personality?
While the internet has lowered the threshold for personal expression, the digital traces it leaves behind do not naturally converge into a complete understanding of an individual.
This problem has been amplified in the era of platform-based internet. While users leave more digital traces, this information increasingly exists within the platform's own ecosystem rather than accumulating continuously around the individual: an article, a post, a video, even with high interaction and dissemination, often exists primarily within the ecosystem of its respective platform. Platforms can record this content, but this information does not naturally coalesce into a personal record.Complete information system.
The first problem brought about by platformization is that it is difficult to establish a unified connection between personal digital traces.An individual may leave behind a wealth of information across different platforms: expressions on short video platforms, opinions on social media platforms, and life records on image platforms, all of which may reveal certain aspects of the individual. However, because this content is scattered across the data spaces of different platforms, it does not automatically establish connections between them, making it difficult to collectively form a continuous personal image.
Besides the difficulty in linking information, platform mechanisms also influence how users generate these digital traces.The platform tends to recommend content that is easy to understand, easy to interact with, and provides quick feedback.This, in turn, shapes how users express themselves.Complex experiences may be compressed into simple techniques, long-term reflections may be broken down into short sentences of viewpoints, and complete experiences may be packaged into more easily disseminated fragments. What ultimately appears on the internet is often only a disseminable aspect, rather than a complete expression of personal experiences, knowledge accumulation, and ways of thinking.
At the same time, the development of content formats has further reinforced this fragmentation.Short videos, images, and short texts all emphasize instant consumption. While this lowers the barrier to expression, it also makes it easier for personal experiences, knowledge accumulation, and thought processes to be broken down into independent fragments. This type of content is very suitable for immediate dissemination, but it is not naturally suited for forming long-term, continuous personal records.
The deeper problem is that the platform organizes information around content, rather than around individuals.The platform focuses on the dissemination and interaction of individual pieces of content, rather than an individual's experiences, long-term interests, or how their views change at different stages. Therefore, even as the amount of information an individual leaves behind increases, it often remains as independent pieces of content, making it difficult to further solidify into a continuous understanding of that person.
This also means that if we want these digital traces to truly become a meaningful long-term accumulation for individuals, we need a medium that can continuously carry content and connect it with others. It must not only record individual pieces of content, but also reflect long-term accumulation, thematic connections, evolution of viewpoints, and practical paths.
Therefore, as AI gradually becomes a new information gateway, a new problem arises:If a person wants future intelligent systems to truly understand them, then in what form should these long-term personal accumulations exist on the internet? This also marks the starting point for rethinking the value of personal websites in the AI era.
3. Personal Blogs: A New Carrier for Personal Digital Assets in the AI Era
If the AI era demands that individuals' long-term accumulation of knowledge be better understood and connected, then the next question to consider is: what kind of information carrier can support this long-term accumulation?
From this perspective, personal blogs have regained a new value—they are not just content publishing platforms in the traditional sense, but information spaces maintained by individuals over a long period of time.
The value of this information space lies not only in its ability to preserve content long-term, but also in its capacity to allow for the continuous accumulation and complete expression of personal content, gradually forming an interconnected information structure. Specifically, this value is primarily reflected in the following three aspects:
- Personal blogs provide an open and stable information space.Unlike the platform-based internet, personal blogs are typically built on independent domains. The content published by the author does not rely on the recommendation mechanisms of any particular platform, but has a long-term, publicly accessible address that can be accessed by search engines, cited by other websites, and continuously become part of the open internet. This openness allows individuals to continuously accumulate content based on their long-term practices. Articles from different times and on different topics are no longer just scattered records on different platforms, but gradually form a continuous accumulation of information.
-
Personal blogs can preserve a more complete expression process.Having established a long-term, stable space for content creation, blogs also offer a more complete narrative space. Compared to platform content that emphasizes immediate dissemination, authors can record how problems arise, how solutions are designed, how practices are implemented, and how their understanding evolves with experience. The value of this long-term record lies in the fact that it showcases not only the final answer, but also a person's way of thinking, practical methods, and cognitive evolution when facing problems.
-
Personal blogs can further organize this informationAs content accumulates, a mature personal blog is not simply a collection of articles. Instead, it connects individual articles into a larger knowledge structure through categorization, series, citations, and associations: articles record specific practices, series present the process of technological evolution, and themes reflect long-term focus. All of these elements together constitute an information system centered around the author.
Therefore, the value of personal blogs in the AI era does not lie in replicating the content dissemination model of the traditional Internet era, nor in competing with various platforms for traffic.
Its more important significance lies in providing individuals with a long-term, open, and stable information storage space, enabling past scattered experiences, practices, and thoughts to continuously accumulate around the same subject and gradually form a more complete information structure.
This structure not only makes it easier for others to understand the author, but also provides a better foundation for future AI to identify, associate, and understand individuals from the public internet.
4. How can we make personal websites easier for AI to understand?
Personal websites can become an important carrier of personal digital personality in the AI era, but this does not mean that the traditional blog structure can directly adapt to the new information environment.
For the past two decades, the development logic of personal websites has primarily been built upon SEO (Search Engine Optimization). The focus of website construction is ensuring content is discoverable by search engines, improving page visibility in search results through keyword matching, page quality, and external links. In this model, the core unit of a website is the "page." For traditional blogs, this means organizing around individual articles, using categories, tags, and search functions to help users locate information. An article that solves a specific problem can potentially gain independent traffic.
However, the advent of the AI era has changed the way website information is understood. As generative AI gradually becomes a new information entry point, websites also need to adapt to this new information environment. This approach is often referred to as GEO (Generative Engine Optimization).
Unlike SEO, which focuses on "how to make pages discoverable by search engines," GEO focuses more on "how to make information on a website understandable by AI." For AI, the value of a website is not just the number of articles it contains, but whether there are clear relationships between these articles and whether they can collectively express some stable informational characteristics.
This also means that websites need to gradually shift their focus from the visibility of individual pages to building a comprehensive content structure.The connections between articles, the evolution between series, the accumulation of long-term themes, and the author information behind the content can all become important bases for AI to understand a website.
To adapt to this change, personal websites need to reorganize their information structure in three aspects.
First, shift from page center to content relationships.
While traditional article categorization and page organization methods still help users browse and retrieve information, they are no longer sufficient to fully represent the information structure of a website. This is because AI needs not only to know what articles are on a website, but also to understand the relationships between these articles.
For example, if a website contains content related to network architecture, server deployment, artificial intelligence, and knowledge management, and these articles are simply arranged according to different themes, AI may only be able to identify a few independent sets of information, but will not be able to determine why they appear together on the same website.
Therefore, a website structure geared towards AI needs to further clarify the relationships between content, including citation relationships, series relationships, thematic evolution relationships, and their connection to the author's long-term practical direction. In this way, the content on the website can gradually form a cohesive information whole with contextual relationships from numerous independent pages.
Second, shift from single articles to long-term accumulation.
In the SEO era, an article that can accurately solve a specific problem can gain independent traffic, so website content often revolves around specific needs.
However, in the AI era, what we need to understand is not just what problem a particular article solves, but also the long-term direction and characteristics that this content reflects. A person's professional abilities, way of thinking, and practical experience usually emerge gradually through long-term recording and continuous expression. Therefore, the value of a personal website lies not only in recording "what problems I have solved," but also in presenting "how I understand and explore these fields."
A website truly possesses the ability to express an individual's long-term value only when content from different times and on different themes can collectively showcase their practical path and evolving thinking.
Third, shift from content collections to author entities
Beyond simply organizing the relationships between content, personal websites also need to more clearly express the author's identity. In the past, users typically got to know the author gradually while reading articles. A reader might enter a website because of a specific question and then learn about the author's background, experience, and characteristics through continued reading. However, in the AI era, AI also needs to determine whether this content comes from the same stable entity and whether there is long-term consistency between these pieces of content.
The "author entity" mentioned here doesn't simply mean adding a personal profile page; it means transforming the author's identity from a byline behind the article into a clearly defined subject within the website's information structure. A connection needs to be established between the author, the article, their professional field, and their practical experience, collectively answering: Who owns this website? What has this person been focusing on long-term? Why can this content form a continuous development path?
Only in this way can the content on the website form a continuous information structure around a clear subject, and gradually reveal the person's long-term practice and cognitive characteristics.
Supplementary explanation regarding "author entity"
The term "entity" here comes from the concept of "Entity" in knowledge engineering.
When AI organizes internet information, it doesn't simply read isolated pages; instead, it attempts to identify different information subjects and establish relationships around these subjects. tangwudi For example, AI needs to determine who the name corresponds to and further associate it with information such as their articles, technical fields, and practical experience.
These relationships can be simply understood as:
- Entity: Represents a specific information subject, such as
tangwudiThis author; - Relationship: Describe the relationship between entities and articles, fields, projects, experiences;
- EvidenceInformation sources used to support these relationships, such as long-running practical articles and technical records.
Therefore, the value of personal websites in the AI era is not just in storing articles, but more importantly, in providing authors with a stable source of information on the public internet and gradually enriching the relationships and evidence surrounding that entity.
For more information on knowledge engineering, please refer to another article: "..."Personal Knowledge Engineering (Part 2): From Cognitive Structure to Knowledge Engineering – Why Does Knowledge Need to Be Structured?》。
5 From Content Recording to Knowledge Structure: Building an AI-Understandable Personal Information System
5.1 From Traditional Blogs to AI-Understanding Websites: An Attempt
After recognizing the changes in information organization that personal websites faced in the GEO era, I began to re-examine my own blog.
For the past few years, I've used this blog as a space to record my personal technical practices: from initially documenting home infrastructure, network configurations, and the use of various tools, to later delving deeper into system architecture, artificial intelligence, and knowledge management. The content on the blog has accumulated along with the development of my personal technical path. This content records my long-term practical process and has reached a certain scale.
However, considering the capabilities required for personal websites in the GEO era, the original blog structure still has room for further optimization: on the one hand, as the number of articles continues to increase, the original connections between different topics are gradually hidden in a large number of independent pages; on the other hand, the technological evolution path accumulated over a long period of time, as well as the author's own information, have not been fully expressed.
Therefore, I began to rethink the blog structure according to the three directions proposed in Chapter 4, and tried to make corresponding adjustments:
- The content, which was originally centered around a single article, gradually forms an information network with internal connections.
-
This allows the articles accumulated over many years to more clearly present personal practical paths and cognitive evolution.
-
This allows the website to not only showcase the content itself, but also to more clearly express the stable entity of the author.
These adjustments are not simply to add blog functionality, but rather to explore a new form of personal website: while retaining the traditional blog recording capabilities, allowing long-accumulated information to form a clearer structure and be more easily understood by both humans and AI.
5.2 From Page to Content Relationship: Building an Internal Information Network for the Blog
The first problem to solve was how to gradually establish connections between content within a blog structure centered around individual articles. The blog's built-in categorization, tagging, and search functions weren't very helpful, so I had to develop new features to try and build relationships between different levels of content.
First, at the article level, I added features such as "Series Article Cards" and "Related Articles" to specific article areas, allowing content with common themes, developmental relationships, or practical backgrounds to be connected. Simultaneously, through the "Right-Side Menu" function, dynamic content recommendations are generated based on semantic association mechanisms, enabling the system to discover potential connections based on the vector similarity between content, thus making the association between articles less reliant on manual classification.
Secondly, I also tried to reorganize the content from the overall website perspective: through a "blog knowledge map" and multiple "specialized knowledge maps", I displayed the main areas covered by the blog and the position of different topics in the overall structure; through the article series menu, I organized content with a continuous evolutionary relationship according to themes, so that scattered articles could form a clearer thematic thread.
The goal of these adjustments is not to add more navigation entry points, but to gradually reveal the previously hidden relationships between content within the blog.
For example, an article about a particular technical practice is no longer just a standalone problem solution, but can connect to relevant background, related content, and broader thematic directions. In this way, individual articles provide specific information, while the relationships formed between articles collectively present a more complete informational framework.
In these ways, blog content gradually transformed from numerous independent pages into a contextually connected information network. This not only makes it easier for readers to understand the relationships between content items but also provides a clearer information foundation for AI to understand the overall structure of the website.
For relevant technical details, please refer to:Build a lightweight knowledge index for your blog“》
5.3 From Single Articles to Long-Term Accumulation: Recording Personal Technological Evolution Path
Unlike content relationships, long-term accumulation is not achieved through a specific function, but is gradually formed through continuous recording and practice.
Fortunately, this isn't something I need to deliberately fuss over for my blog. In fact, from its inception, my blog wasn't a collection of content built around a single field, but rather it expanded continuously with my personal technical practice—from early home infrastructure, network architecture, and server deployment, to later Cloudflare, system architecture, artificial intelligence, and knowledge management, recording the process of how the direction of technical exploration evolved with practice.
These contents superficially involve multiple technical topics, but they are not entirely independent: each change in technical direction comes from the accumulation of problems in previous practices, and the problems encountered in the previous stage often become the starting point for subsequent technical explorations. Early exploration of infrastructure promoted subsequent thinking on system design, website architecture and information organization methods, and these experiences further became the foundation for exploring knowledge management and personal digital expression in the AI era.
From this perspective, the value of my blog lies not only in recording "what problems I've solved," but also in presenting "how I understood and explored these problems." When content from different stages collectively reflects an individual's technical experience and way of thinking, the long-term accumulation itself becomes a significant value that distinguishes a personal website from ordinary content platforms.
5.4 From Content Aggregators to Author Entities: Enabling Websites to Express "Who Creates This Content"“
5.4.1 From Author Information to Author Entity
In traditional blogs, the author is usually just a piece of information accompanying the article. For example, an article might display "Author: tangwudi" at the bottom, which is sufficient for human readers. Readers can infer from the website context that this is the blog maintainer, someone who regularly publishes technical articles, and that this name may also have corresponding accounts on other platforms.
However, for a machine, "tangwudi" is first and foremost just a string—it cannot determine from this name alone whether it is a real person, a user account, an organization name, or a brand name. Similarly, it cannot automatically confirm whether "tangwudi" in the blog also corresponds to a corresponding account on GitHub or other platforms.
This is precisely the difference between how humans and machines understand information: humans can establish identity judgments through context, while AI requires more explicit information descriptions. Therefore, in a website structure that AI can understand, the author cannot simply be a name in an article, but needs to be a clearly described information object, that is, the author entity.
The "entity" mentioned here is not simply adding a personal profile page or adding more text descriptions to the author. Rather, it is about establishing a stable identity for the author in a structured way and linking information such as articles, professional fields, and practical experience on the website to this identity.
For example, my blog website can use structured data to store my username. tangwudi Described as a Person Entities of type:
{ "@type": "Person", "@id": "https://blog.tangwudi.com/#person", "name": "tangwudi" }
in,@type Used to tell the machine: This object belongs to Person Type, which represents a real person;name The code name used to describe this person is tangwudi;and @id This is used to provide a unique identifier for this entity:”https://blog.tangwudi.com/#person".
The string herehttps://blog.tangwudi.com/#personThis isn't a URL requesting access to a webpage, but rather more like a unique ID in a database. It indicates that within the website's information structure, there exists a Person entity representing the author, tangwudi. Subsequently, if other content on the website references this identifier, the machine can recognize that they all point to the same entity. For example, an article might contain:
{ "author": { "@id": "https://blog.tangwudi.com/#person" } }
Another article also includes:
{ "author": { "@id": "https://blog.tangwudi.com/#person" } }
From the AI's perspective, it sees that both articles are associated with the same Person entity. This means that the author's identity behind the articles is no longer just repeated text, but has become a stable information node.
It should be noted that@id The specific string following the name doesn't have a fixed format requirement, as long as it uniquely identifies the entity and remains consistent within the website. In fact, website maintainers can define different formats according to their needs. @idFor example, using custom numbers or other identifiers with clear meaning.
The reason for using here https://blog.tangwudi.com/#personThis is mainly because the WordPress SEO plugin Rank Math generates this entity identifier by default based on the website domain. Since this identifier already meets the requirements, I didn't redesign it during the actual modification, but directly used the default form generated by Rank Math.
In addition, author entities can also be accessed through... sameAs This involves establishing a connection between identities on other platforms, such as linking my GitHub account.
""sameAs": [ "https://github.com/tangwudi1979" ]
It doesn't mean "These are my other links," but rather it describes the content of the blog. tangwudi With GitHub tangwudi1979 They all point to the same real-world entity. In this way, AI can attribute information distributed across different platforms to the same author entity.
In summary, for personal websites, the significance of author entities lies in ensuring that the content on the website can be clearly attributed to a specific, stable individual entity—in simpler terms, author entities can improve the recognizability of an individual's identity on the internet.
For example, my blog can provide machines with clearer clues about my identity, allowing... tangwudi This name makes it easier to establish a stable association with my blog, GitHub account, and the technical content I've accumulated over time. As these associations continue to build, AI... tangwudi The understanding of it will no longer be just an ordinary string, but will gradually form a clearer sense of identity.
In layman's terms, it means continuously supplementing the contextual information about the name so that the AI can more easily "recognize" the specific entity behind the name, thereby reducing confusion with other accounts, nicknames, or other meanings with the same name.
This is also the core meaning of personal websites evolving from "content collections" to "author entities".
5.4.2 Enabling Personal Websites to Have Clear Author Identity Through JSON-LD
In the previous section, we discussed the structured data used to represent author entities:
{ "@type": "Person", "@id": "https://blog.tangwudi.com/#person", "name": "tangwudi", "sameAs": [ "https://github.com/tangwudi1979" ] }
This type of structured data is typically used JSON-LD (JavaScript Object Notation for Linked Data) The format representation—it is not the content displayed to users on a webpage, but a machine-understandable description provided to search engines and AI systems to explain what the entities on the webpage are and what relationships exist between them.
Therefore, to give a personal website a basic author entity, it is essentially about adding JSON-LD structured data describing that person to the webpage.
For static websites, this can usually be done directly in the HTML template. <head> Add the following code to the section:
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Person",
"@id": "https://blog.tangwudi.com/#person",
"name": "tangwudi"
}
</script>
In this way, when search engines or AI crawlers access web pages, they can read the author entity information within them.
However, for dynamic websites like WordPress, where page content is typically generated dynamically by the theme, plugins, and database, manually modifying each page is not suitable. A more reasonable approach is to utilize existing SEO plugins to generate the basic structure and then adjust it according to actual needs.
For example, the SEO plugin I use for my blog is Rank Math (for more information, please refer to: [link to relevant documentation]).Rank Math SEO Setup and Optimization: From Installation to Introduction to Commonly Used Functional ModulesWhile it already supports Schema.org structured data generation and can automatically output relevant entities such as websites, articles, and authors, the problem is that the data structures generated by the plugin are mainly geared towards general SEO scenarios and may not necessarily meet the needs of building a personal digital identity. For example, in my case, the author identity and website entity identity were not unified into the same entity in the initial schema generated by Rank Math.
At the website level, a default subject entity will be generated:
{ "@type": ["Person","Organization"], "@id": "https://blog.tangwudi.com/#person", "name": "tangwudi" }
Additionally, Rank Math will generate author and publisher associations within the article:
{ "author": { "@id": "https://blog.tangwudi.com/me/tangwudi/#author" }, "publisher": { "@id": "https://blog.tangwudi.com/#person" } }
In other words, there are actually two different identity references on the same website:author Pointing to the WordPress author's identity;publisher It points to the identity of the website's main entity.
Although they all used names tangwudiHowever, due to @id Despite the differences, they remain two distinct entities in the schema semantics. This is a common problem encountered by dynamic websites when building personal digital identities: plugins can generate structured data, but they don't know how the website operator wants to define their long-term identity relationships.
Therefore, instead of continuing to use the author body generated by Rank Math by default, I used the one provided by Rank Math via PHP:
rank_math/json_ld
The filter performs secondary processing on the final output JSON-LD, mainly making three adjustments:
First, the main entity of the website is clearly defined as Person.
Rank Math (default):
{ "@type": ["Person","Organization"], "@id": "https://blog.tangwudi.com/#person" }
Revised to:
{ "@type": "Person", "@id": "https://blog.tangwudi.com/#person", "name": "tangwudi" }
The reason is simple: this website is essentially a personal technical blog and does not have a corporate entity that needs to be identified, so there is no need to retain the Organization type.
so,https://blog.tangwudi.com/#person It thus becomes the core entity that represents the author's identity as the sole representative of the entire website.
Second, the authors of the articles are all attributed to the core Person entity.
By default, Rank Math articles may reference entities on the author's archive page:
{ "author": { "@id": "https://blog.tangwudi.com/me/tangwudi/#author" } }
This association points to the author's archive page, not a personal identity entity that is intended to be accumulated long-term. Therefore, the author citation in the article should be modified as follows:
{ "author": { "@id": "https://blog.tangwudi.com/#person", "name": "tangwudi" } }
In this way, all articles on the website will be clearly associated with the same author entity.
Third, add an identity attribute to the Person entity.
The basic Person entity only states that "this person exists," but it doesn't specify this person's domain or relationships. Therefore, I further added:
{ "description": "A technology practitioner exploring infrastructure, AI, and knowledge engineering through systematic thinking and engineering practice.", "sameAs": [ "https://github.com/tangwudi1979" ], "knowsAbout": [ "Infrastructure Engineering", "Web Architecture", "Artificial Intelligence", "Knowledge Engineering" ] }
in:description Used to describe who this entity is;sameAs Used to declare that accounts on other platforms on the Internet belong to the same person;knowsAbout This describes the knowledge domain that the entity focuses on and practices over a long period of time.
Ultimately, through these adjustments, the structured data on the webpage is no longer: This article was published by an author named tangwudi. Instead, it becomes: There exists a Person entity on the internet named tangwudi, with a clear website identity, a GitHub connection, and a continuous output of content related to infrastructure, AI, and knowledge engineering.
However, it's important to note that establishing the basic author entity is only the first step. JSON-LD solves the problem of letting machines know that "this website has a clearly defined author." But it cannot fully describe: which areas this author focuses on in the long term; what technical experience they have; or why these articles form a continuous development path.
in other words:The author entity addresses the issue of "identity assertion," not the issue of "personal knowledge representation."
Google Rich Text Detection and Verification
After completing the schema adjustments, you can verify the final output using Google's rich media search tools. The results show that Google is correctly resolving the relationship between the article and author entities, and you can also see the identity attributes I added to the author entity:


5.4.3 Enhancing the Author Entity: Enabling AI to Understand "Who is this person?"“
After establishing the basic author entity, the website can clearly tell the machine that the content belongs to the same author. However, for AI, simply knowing "who this person is" is not enough.
For example, the basic author entity can indicate that the author of this website is... tangwudiThis author is associated with a GitHub account; all articles on the website are published by this author. However, it doesn't further explain: which fields this author focuses on in the long term; what practical experience they possess; why this content appears on the same website repeatedly; or what professional direction and cognitive characteristics this content reflects.
Therefore, beyond the basic author entity, it's necessary to supplement it with information that describes the author's characteristics, allowing the author entity to gradually expand from a simple identifier into a more complete information node. Based on this idea, I attempted to add some information entry points for future AI environments, primarily by adding... author-profile.json and llms.txt These two files provide more context for machines to understand author identity and website content.
use author-profile.json Describe author information
To supplement the author entity information, I created author-profile.json The document is used to centrally describe the author's basic information, technical field, long-term practice direction, and content characteristics.
With JSON-LD Person Different entitiesauthor-profile.json It is not used to declare "this person exists", but to further describe: who this person is, and in what aspects this person's long-term accumulation is reflected.
For example, a simple description of author information could look like this:
{ "name": "tangwudi", "description": "Personal technical blogger, chronicling practical experiences in home infrastructure, network architecture, AI, and knowledge management.", "fields": [ "HomeLab", "Network Architecture", "Cloudflare", "Artificial Intelligence", "Knowledge Engineering" ], "focus": [ "Technical Practice Recording", "System Architecture Exploration", "Personal Knowledge Management" ] }
The focus here is not on the fields themselves, but on expressing information that is originally hidden in human text through a structured approach.
In this way, when AI understands website content, the information it obtains is no longer just: a name tangwudi The author has published several articles. It can also be understood as a technical author who has long practiced and continuously documented their experience in infrastructure, network architecture, artificial intelligence, and knowledge management.
In actual deployment, the files can be placed directly in a publicly accessible location on the website, for example:
https://blog.tangwudi.com/author-profile.json
These documents can serve as machine-oriented information entry points, providing a more structured context for future search systems, AI agents, or other automated tools that may access the website.
In this way, a more complete information relationship is formed between the author entity, website content, and author information description.
use llms.txt Provides an AI-guided information portal
In addition to describing the author, you can also use site documentation for AI systems.llms.txtThis provides a more direct website context. For example:
https://blog.tangwudi.com/llms.txt
The following is the content of my blog file llms.txt:
# tangwudi blog This website is the personal knowledge base of tangwudi, a technology practitioner exploring how complex technical systems are built, operated, understood, and transformed into structured knowledge. ## Author Author profile: [https://blog.tangwudi.com/author-profile.json](https://blog.tangwudi.com/author-profile.json) tangwudi focuses on: - infrastructure engineering - web architecture - artificial intelligence - knowledge engineering The author's methodology includes: - first principles thinking - system evolution analysis - engineering practice - knowledge structuring ## Core Topics ### Infrastructure Engineering Exploration of: - home data centers - self-hosted services - networking - virtualization - Linux and container technologies ### Web Architecture Research and practice around: - WordPress architecture - Cloudflare ecosystem - CDN and security optimization - reliable personal websites ### Artificial Intelligence Exploration of: - large language models - embeddings - semantic search - AI-assisted workflows ### Knowledge Engineering Research into: - knowledge organization - semantic indexing - structured knowledge representation - AI-readable knowledge systems ## Knowledge Map The [Blog Knowledge Map](https://blog.tangwudi.com/roadmap) provides a structured overview of the site's core knowledge domains, topics, and related articles. It can be used as the primary navigation entry for understanding the structure and relationships of the site's content. ## Content Philosophy The website focuses on understanding technologies through their underlying principles, historical evolution, practical implementation, and architectural trade-offs. The goal is not only to document solutions, but to transform engineering experience into reusable knowledge structures.
Compared to regular navigation on a webpage,llms.txt It provides a website summary geared towards AI understanding. It doesn't replace website content, but rather helps AI quickly understand: what the website is; what the author focuses on; and which content is representative.
However, it should be noted that,at present llms.txt There is no de facto standard yet, and different AI systems support it to varying degrees.This article views it merely as an attempt at information expression in a future AI environment, rather than a mature, universal standard.
With the addition of author information, structured data, and these additional enhancements, the author entity is no longer just a name and identity identifier, but gradually becomes an information node that includes identity, experience, professional direction, and long-term accumulation.
At a deeper level, the goal of these initiatives is not simply to add an author bio, but to continuously enhance the open internet. tangwudi This relates the entity to its associated content. Through continuously accumulating practical articles, technical records, and structured information, AI can more accurately distinguish information from different sources and gradually establish a stable understanding of the entity represented by the name—in simpler terms, the hope is that in the future, when the name is mentioned… tangwudiAI can prioritize thinking of me, an individual subject who has been recording and practicing for a long time.
6 When AI begins to understand a personal website
6.1 AI-generated author profiles based on publicly available information
The previous chapters discussed the capabilities that personal websites need in the AI era, and how I can make the content, relationships, and author identities in my blog easier for machines to understand through structural transformation.
However, the final effect of these transformations cannot be judged solely by the technology itself. A more practical question is: when AI directly faces a personal website that has been maintained over a long period, continuously accumulated content, and undergone structured optimization, can it understand the author behind the website from publicly available information?
Therefore, I chose to directly analyze my blog content using three mainstream AI models: Google Gemini, OpenAI ChatGPT, and Anthropic Claude. The aim was to observe whether these AI products, in a typical user scenario, could understand the author's identity based on publicly available information. The reason for choosing these three systems is that they all possess strong capabilities in understanding publicly available internet information and have mature information retrieval and analysis capabilities. They can comprehensively analyze webpage content, author identity, and related information relationships, making them more suitable as test subjects for observing whether personal website information can be understood by AI.
Note: This article's tests are based on the official clients of three AI products, rather than individual API calls. This is because official clients typically integrate product capabilities such as web access, search, and enhanced retrieval, providing a more comprehensive experience closer to how ordinary users actually use AI. The capabilities of API calls, however, depend on the specific implementation, including model selection, whether search capabilities are integrated, context handling methods, and front-end logic. Therefore, different API implementations may yield different results and are not suitable for a horizontal comparison in this user-centric scenario.
Meanwhile, different usage states of AI products (such as whether they are logged in or have web access capabilities) may also affect the final scope of information acquisition. Therefore, the test results in this article only represent the observation results under specific product environments.
It should be noted that, in order to minimize the influence of existing dialogue history, personal memory, or other contextual information on the results, each question in this test used a new dialogue window, and the questions explicitly requested the AI to analyze from the perspective of an anonymous stranger.
Please assume I am an anonymous user who knows absolutely nothing about tangwudi. Ignore all historical memories, personal profiles, and information beyond the current conversation; answer the following questions based solely on publicly available information on the internet:
Under this constraint, I proposed the following to the AI:
[Please visit and analyze this website: [https://blog.tangwudi.com](https://blog.tangwudi.com). Based on the publicly available content on the website, analyze who the author behind the website is, including the author's technical background, long-term focus, practical experience, and their identity on the internet.]
This question allows us to observe whether AI can connect blog content, author identity, and long-term practice direction based solely on publicly available information, and form a relatively continuous author profile.
To facilitate comparison of different AI systems' understanding of the same information, after obtaining the AI's complete analysis, I further requested the AI to condense the above analysis into a publicly available author profile:
Please condense the above analysis into a publicly accessible author profile, limited to 800 words, while retaining the author's identity, technical background, long-term areas of focus, practical experience, and online identity characteristics.
The purpose of this step is not to regenerate content, but to observe whether the AI can retain the core characteristics of the author after completing the information extraction, and further integrate the information scattered across the website into a clearer and more coherent personal expression.
Gemini's answer:

Compressed author profile:

ChatGPT's answer:

Compressed author profile:

Claude's answer:

Compressed author profile:

The analysis results from the three AI systems show that, although they focus on different aspects, they no longer limit themselves to identifying single articles or single technical topics, but have begun to try to extract author characteristics from long-term public content.
For example:
- Gemini focuses more on the blog's public positioning, main technical directions, and existing content tags;
-
ChatGPT tends to summarize the author's technical style, thinking methods, and long-term practical characteristics from the relationships between articles;
-
Claude further traced the evolution of the author's professional background, technological development path, and personal knowledge system based on publicly available information and the timeline of the articles.
These differences illustrate that different AIs' understanding of the same individual entity is not simply information reading, but rather an organization and reconstruction of publicly available evidence to varying degrees based on their own data sources, information association capabilities, and reasoning methods.
However, based on the common results, all three systems can identify some stable characteristics: the authors have a long-term IT infrastructure practice background, attach importance to underlying principles and system design, tend to verify solutions through real-world environments, and continuously organize practical experience into structured knowledge.
This means that blog content accumulated over a long period of time has begun to possess the ability to express the author's identity, professional focus, and cognitive characteristics.
6.2 Entity Association Status of the Name “tangwudi”
Beyond analyzing blog author profiles, I also made a more direct observation: when AI faces... tangwudi When using this name, it is important to determine whether it can be stably associated with the same individual entity.
To avoid the influence of factors such as historical conversations and user profiles on the test results, the test will still be conducted from the perspective of an anonymous stranger.
Please assume I am an anonymous stranger who knows absolutely nothing about tangwudi. Ignore all historical memories, personal profiles, and background context regarding me, and answer this question solely based on publicly available information on the internet: "Who is tangwudi?"“
Gemini's answer:

Ghatgpt's answer:


Claude's answer:

It's important to note that different AI systems may not have entirely consistent understandings of the entities formed from the same publicly available information. For example, Gemini's response focused more on the blog's public positioning, main technical directions, and existing content tags; ChatGPT further combined the relationships between articles to summarize the author's technical style, practices, and long-term thinking characteristics; Claude tended to trace the author's development path from their public experience, timeline, and website ecosystem.
This difference does not mean that a certain result is necessarily correct, but rather reflects the differences between current AI systems in terms of information sources, data updates, content association capabilities, and reasoning methods. The information left by the same person on the Internet needs to go through the information processing processes of different AI systems to form the final entity understanding result.
Based on the current test results,tangwudi This name has already established a relatively clear association with blog.tangwudi.com and related publicly available technical content across several mainstream models. Although different AI approaches have different focuses, it is generally possible to identify from publicly available information an individual entity that maintains a personal technical blog, continuously records IT practices, and explores areas such as home data centers, network architecture, AI, and knowledge systems.
However, this correlation has not yet reached a level that all AI systems can reliably and accurately identify. For example, in the testing of some models (such as Qwen and DeepSeek), directly asking "“tangwudi When identifying "who it is", the recognition results are not stable. However, by adding limiting conditions such as "blogger", it is easier to obtain accurate association results.
This may be related to the differences in the information sources and information processing methods that different AI systems rely on. Some models also mentioned in their responses that their understanding of public information mainly relies on Chinese Internet content. Therefore, their recognition effect on public information from Chinese personal websites and the local Internet ecosystem may differ from that of other large models that face the global Internet.
Of course, I also have a slightly "whimsical" long-term goal: I hope that one day in the future, when I directly ask "Who is tangwudi?" in different AI systems, they can all reliably link it to me as an individual based on publicly available information and form a relatively consistent and complete understanding.
If this can truly be achieved, then perhaps the transformation from a username to a personal entity with a stable identity in the open internet, capable of being continuously understood by AI, will be truly complete.
The testing methodology in this section is still not rigorous enough. Although a new dialogue was used during the test and historical memory was explicitly ignored, personalized factors at the account and device levels can still affect how the AI acquires and processes information, such as region, language preferences, and past usage behavior—these factors cannot be completely eliminated simply by the instruction to "ignore historical memory."
Therefore, if rigorous verification is required tangwudi Ideally, the public internet connection to this blog should use a third-party device and account completely unrelated to me, and keep the default settings so that a real stranger can directly ask the mainstream AI, "Who is tangwudi?" This way, the results will be closer to test results without the interference of personal history.
Therefore, what this experiment proves is that after these structural modifications, AI can indeed... tangwudi This name is associated with this blog; however, based solely on this set of experiments, it is impossible to determine how much of this association comes from structured data and how much may be influenced by personalized factors. To further separate these two factors, more rigorous controlled experiments are needed, which are currently impossible to conduct due to personal limitations and can only be verified in the future.
7 From Digital Personality to Digital Twin: The Future Evolution of Personal Knowledge Assets
In the traditional internet era, a person's digital presence was primarily manifested in accounts, articles, works, and publicly available information. This content could prove "what this person left behind," but it was difficult to fully present "how this person thought." In the AI era, however, acquiring knowledge, organizing information, and generating standard answers are becoming increasingly easy. What is truly difficult to replicate are a person's judgment methods, accumulated experience, and problem-solving paths developed in real-world practice.
For personal websites, the modifications mentioned earlier primarily address a practical problem: how to present the information accumulated over time in a more complete and coherent manner in the AI era. However, from a longer-term perspective, the significance of these improvements goes beyond simply optimizing a website; it's about gradually building a personal identity.Digital assetsThe real value of these digital assets lies not only in the content a person leaves behind, but also in the experience, methods, and ways of thinking that gradually develop behind that content.
As these long-term accumulations gradually form connections and become understandable to AI, they begin to transform from mere information records into part of an individual's digital personality. Furthermore, as AI's understanding of individuals develops, this digital personality may become a crucial foundation for future personal digital twins.
Supplementary Notes on "Personal Digital Twins"
“The term "digital twin" was initially applied primarily to fields such as industrial manufacturing and urban management. It refers to mapping real-world objects through digital models and continuously updating them based on real data.
When this concept is extended to the personal realm, the focus is not on simply replicating a person's digital image, but on using long-term accumulated digital information to continuously form a dynamic mapping of that person.
When this mapping is complete enough, AI may be able to further simulate a person's way of thinking and judgment logic, thereby exhibiting similar analytical and decision-making tendencies when facing specific problems.
For example: Why did this person ultimately choose a certain option? Why did they abandon several other seemingly better options? How do past experiences influence today's judgment? These are the kinds of things that truly belong to an individual's cognitive abilities.
This ability cannot be acquired through short-term data collection, nor can it be generated simply through a single model training session. It requires authentic recording, continuous accumulation, and the constant transformation of personal experiences into understandable information.
Personal blogs are perfect for preserving such long-term records.An article records a specific problem and its solution process. Over time, these records will gradually reveal a person's practical experience, choices, problem-solving methods, and cognitive changes.When this information can be continuously understood and correlated by AI, it will no longer be just a personal historical record, but may become a new personal advantage.
In the future, a new group with digital advantages may emerge: they may not be public figures in the traditional sense, nor do they necessarily have a large following or social influence, but because they continuously leave behind authentic, continuous, and valuable information records on the open internet, AI can associate relevant information with a stable individual entity, such as directly knowing "who tangwudi is." Of course, this association also has a very realistic prerequisite:The more unique the name, the easier it is to eliminate ambiguity.— "tangwudi" is definitely much easier than "Zhang San" or "Li Si".
In this sense, these people can be called "digital asset accumulators" in the AI era.The long-term value of personal websites lies in providing an open, stable, and continuously traceable platform for this kind of accumulation.