Home / Transforming a corpus of 7,000 pages into living knowledge

Transforming a corpus of 7,000 pages into living knowledge

Olivier Carrère 9 min read
View as Markdown
On this page

From 7,000 pages to daily insights

How can a non-profit turn 7,000 pages of articles—a vast sea of text—into content that is alive, readable, and meaningful every day?

1,800
Oral teachings
1.8M
Words (1999–present)
7,000
Pages in archive
This project extends a broader effort to leverage a massive corpus into a discovery system that makes exploration smarter, more meaningful, and deeply human. Instead of letting a monumental archive gather digital dust, the goal is to let it speak again
one insight, one quote, one topic at a time. The technical workflow behind building this corpus is described in converting Word files to SEO-optimized web pages with AI.

The transformation from dormant files to an active editorial system follows a clear pipeline:

  1. 1. Raw corpus: 1,800 articles / 7,000 pages
  2. 2. Clean & normalize Markdown
  3. 3. Enrich: topics, tags & semantic scores
  4. 4. Discover: topic mapping & query matching
  5. 5. Recompose: digests, daily articles & quotes
  6. 6. Engage: web, print & micro-blogging
  7. 7. Living archive: ongoing reader feedback

Closing the audience gap with AI-curated content

Analytics from the non-profit’s website show a clear pattern: the audience is predominantly aging and male.

While loyal and deeply engaged, this profile highlights a challenge for the organization’s future
its reach remains limited, leaving out younger visitors and potential readers whose interests, language, and expectations differ.

To ensure greater inclusivity and enable a renewal of generations, the non-profit’s communication now aims to better engage women and people aged 25 to 35, fostering a more balanced and sustainable group over time.

1. What we know

Observed data

Analytics reveal an audience that is predominantly aging (59.6% over 55) and male (77.8%). The reach is loyal but structurally limited.

2. What we explore

Audience questions

Public forums, discussion boards, and exploratory AI queries surface what 25–35 year-olds actually worry about: time, authenticity, and safety.

3. What the corpus provides

Editorial bridge

1,800 historical articles already contain answers to those modern questions. The task is retrieval, curation, and accessible recomposition.

Supporting data, not full diagnosis: The demographic charts show who currently reaches the site. They identify an audience-renewal challenge, but they cannot tell us what prospective readers need—that requires exploring real questions.

Age Distribution of Website AudienceAge Distribution of Website Audience18–2425–3435–4445–5455–6465+302826242220181614121086420Percentage (%)
Figure 1 — Observed age distribution: 59.6% of website visitors are 55 or older, with only 3.7% aged 25–34.

To bridge this gap, the corpus of 1,800 articles is being used not only as a knowledge base but also as a dialogue tool. By aligning the content’s depth with the audience’s real concerns, the project aims to make long-form wisdom resonate with new generations of readers.

Gender Distribution of Website AudienceGender Distribution of Website AudienceFemaleMale1009080706050403020100Percentage (%)
Figure 2 — Observed gender distribution: 77.8% male vs. 22.2% female audience proportion.
As a first step, we asked an automated system to scan public forums, discussion boards, and thematic websites, gathering the ten most common questions this audience expresses about the non-profit’s field of activity. These questions become entry points
bridges between lived curiosity and archival insight.
A suspension bridge spanning a wide gorge
Questions gathered from public discussions act as bridges between modern lived curiosity and historical archival insight.

The bridge workflow links reader questions directly to the corpus:

  1. Audience question
  2. Identify topic
  3. Retrieve corpus material
  4. Curate reading path
  5. Measure engagement

This radar chart shows the website audience across six age groups, alongside gender distribution. Each axis displays the Website Audience ratio relative to the general population and the corresponding male and female percentages. Values above 100% indicate that the website has proportionally more visitors than the general population in that age group.

Website Audience and Gender Distribution by Age Group18–24 (4%)25–34 (60%)35–44 (104%)45–54 (156%)55–64 (236%)65+ (284%)MaleAgeFemaleWebsite Audience and Gender Distribution by Age Group
Figure 3 — Audience renewal gap: website demographics indexed against general population across age brackets.
Through this approach, the archive evolves from a static collection into a responsive, audience-aware ecosystem
one that listens as much as it speaks.

Exploring audience concerns through Q/A sessions

To better understand the interests and hesitations of different audience segments, we ran exploratory Q/A sessions with ChatGPT. These sessions helped surface recurring questions from readers, which guide content curation and engagement strategies.

For example, when querying:

Example exploratory query: “What are 25- to 35-year-olds most concerned about when it comes to mindfulness?”

The responses highlighted three primary areas of hesitation:

1. Time & consistency

Concern

Many feel their schedule is too crowded or erratic to sustain regular practice, creating guilt or abandonment before habits take root.

2. Authenticity

Concern

Skepticism about commercialization, trendy app-based wellness marketing, and shallow interpretations stripped of depth.

3. Effectiveness & safety

Concern

Apprehension that practice might be ineffective, or worry that silent introspection could inadvertently unearth difficult emotions.

The non-profit operates with a very limited budget and cannot afford formal commercial market research. To make well-informed decisions, it relies on ChatGPT as a cost-effective research tool for exploring audience concerns, identifying trends, and guiding content strategy based on publicly available information.

To keep the system balanced and transparent, automation is applied across four distinct roles:

Audience exploration

Public web

Scan public forums and thematic discussions to surface common questions, hesitations, and recurring concerns without expensive consulting studies.

Corpus analysis

Metadata & tags

Parse 1,800 Markdown files to assign domain tags, semantic scores, and cross-references across the collection.

Contextual retrieval

Search & match

Query the scored archive to match identified audience questions with relevant historical articles and specific quotes.

Editorial recomposition

Multi-format

Assist human editors in assembling daily articles, 200-page thematic digests, and micro-blogging quotes from the same source files.

Human-in-the-loop principle: Automation serves as an instrument for attention, not distraction. AI helps triage a vast archive that no single editor could re-read in full; human curators decide which insights are genuine, faithful, and worth publishing.


Curating content and quotes from identified topics

Once the system has identified key topics within the corpus, the next step is to leverage those insights for discovery and engagement.

  1. 1. Identify topic
  2. 2. Query corpus
  3. 3. Curate articles
  4. 4. Extract quotes
  5. 5. Publish daily flow

For each topic, the workflow involves:

  • Querying the corpus to surface the most relevant articles.
  • Curating these articles to make them more visible on the website and to compile them into themed digests for deeper reading.
  • Extracting quotes from these articles that resonate strongly with each topic, ready to be shared daily on micro-blogging platforms.
Corpus Topic Curation Pipeline Diagram

1.8 million words
1,800 articles
7,000 pages

1 article per day
for 5 years
on website

Quote ranking
& retrieval

Print - 200 pages
digests by topic

Micro-blogging

Topic 1
Time and consistency

Topic 2
Authenticity

Topic 3
Effectiveness & safety

Figure 4 — Branching editorial streams: a single 1,800-article corpus powers daily web publishing, social quote extraction, and thematic print booklets.

This process transforms the archive from a static repository into a living, thematic ecosystem, where content is both discoverable and shareable. Readers can explore topics in depth, while daily quotes keep the conversation active and ongoing, bridging the gap between long-form articles and real-time engagement.


From files to flow

The foundation is a structured corpus of ~1,800 oral teachings (1.8 million words spanning from 1999 to present) in .mdx files with YAML frontmatter: the single source of truth.

Before any automated processing can help, the source data must be clean and consistent
duplicates removed, formats standardized, and terminology aligned. This ensures a high-quality foundation for metadata generation, indexing, and content enrichment:
  1. Clean Markdown (.mdx) source
  2. Normalize formats & terminology
  3. Generate structured metadata
  4. Branch into web, print & quotes

From that clean source, three parallel publishing streams operate:

Daily web publishing

Website

Continuous daily publishing on zen-deshimaru.com (one article per day), sustaining an ongoing rhythm for over five years.

Thematic print digests

Booklets

Over 20 seasonal A4 two-column booklets (200 pages each) compiling focused selections for in-person retreats.

Master index volume

400 pages

A comprehensive master reference combining index-terms.json occurrences and AI-generated definitions from index-glossary.json.


Scoring, metadata, and thematic discovery

Metadata transforms the corpus into a discoverable and analyzable knowledge base. Automated workflows assist in reading each article, identifying topics, tags, and relevance scores, and maintaining consistency across the entire collection. For a detailed look at how AI scores and tags individual Markdown articles, see transforming meditation class transcripts into an AI-powered discovery service.

The metadata layer performs three distinct functions:

  • Topic categorization: automatically labeling content by domain or focus.
  • Scoring and ranking: assigning interest scores to highlight resonance and contextual depth.
  • Content enrichment: adding summaries, cross-references, and definitions for technique-related terms.

A Python-based workflow evaluates every article, assigning semantic scores that reflect each topic’s presence and intensity. The result is a harmonized dataset, where every page carries structured metadata, ready for discovery, filtering, or print curation.

Triage, not truth: An interest_score of 10 is a model’s judgment, produced in the confident register models always use: it isn’t a measurement, and it has had no contact with an actual reader’s interest. The same holds for the “ten most common questions” scanned from forums and the resonance ranking on quotes: useful heuristics for where to look first, not evidence of what the audience actually wants. They’re a way to triage 7,000 pages no human could read in full, which is genuinely valuable—but triage, not truth. Let the model sort the pile; let a person who knows the material decide what rises from it.


From AI-scored digests to thematic discovery dimensions

Once scored and organized, the corpus is ready to be recomposed into focused digests: each a lens for re-seeing the same field of knowledge from new angles.

One source, multiple dimensions: The same 1,800 articles can be sliced by pedagogical topic, historical chronology, or audience dilemma. Each thematic digest is a different lens on a single coherent archive, creating tailored reading paths without fragmenting the underlying data.

These digests reveal different dimensions of the corpus, guiding readers to explore content thoughtfully and meaningfully.


From quotes to conversations

The best quotes don’t stay confined to pages: they move. Through micro-blogging, daily publishing, and social curation, the project opens channels for dialogue: the corpus becomes a conversation.

  1. Long-form archive
  2. Quote extraction
  3. Resonance ranking
  4. Daily micro-blogging
  5. Reader conversation
  6. Refined search & curation
Each quote, ranked by resonance and context, is shared not as static text but as a living signal
bridging the long-form archive and the fast-moving web. Automated systems support this process by identifying patterns, key sentences, and recurring motifs, ensuring the most meaningful insights circulate outward.

Toward a living, AI-powered knowledge archive

This is more than content recycling; it’s content renewal. A living editorial process transforms static archives into intelligent, evolving knowledge.

Readers can:

  • Discover articles by topic, relevance, or context.
  • Explore cross-linked insights through search or guidance.
  • Read beautifully designed print anthologies: curated 200-page digests generated from the same dataset.

The goal remains human: to distill clarity, preserve depth, and foster presence. Automation serves as an instrument for attention, not distraction: a tool to bring the archive to life again.

What emerges when archives breathe again isn’t noise: it’s continuity. A flow of attention, insight, and care that transforms reading into renewal.

Summing up

External sources

Hero image: “Library & Reading Room — Rand Club” by Andrew Moore, licensed under CC BY-SA 2.0.

Follow Olivier Carrère on LinkedIn

Continuous writing on docs-as-code, DITA XML, YAML, and AI-assisted documentation pipelines.

Follow ↗