MADKIDS Gazette · Engineering case study

Preserving hundreds of stories
without losing their identity.

A family archive built to preserve my children’s COVID-era newsletters, original work, and more than 300 articles.

MADKIDS Gazette logo
MADKIDS GAZETTEA family-run newsletter for curious kids

From a family learning project to a durable archive

During COVID, my children created newsletters to continue their education and explore extracurricular interests. After they had created more than 300 articles, I wanted to preserve their work in a website that made it easy to browse, rediscover, and share.

The difficult part was organizing the material. Each article needed a category, a durable identity, and a reliable relationship to its original newsletter. A migration that merely copied text would lose author attribution, volume relationships, imagery, and the context in which the children created the work.

The actual framework and delivery model

AstroReact islandsJavaScript / TypeScriptJSON contentNode.jsCloudflare static assetsGoogle SlidesSharpPlaywrightaxe-core

MADKIDS uses Astro static generation, with React islands only where an interactive directory needs real-time filtering. The repository has no runtime database, account system, or application API. Publication metadata and content are stored in version-controlled JSON and Astro files, with media under public assets.

Astro generates story pages, topic pages, author collections, interactive presentation routes, and the XML sitemap. Cloudflare serves the generated files. Google Slides embeds preserve interactive editions alongside article records and locally stored thumbnails.

Separating imported records from the canonical catalog

The source import dataset retains extracted content and original metadata. A separate canonical catalog supplies the public home page, archive, topics, authors, related stories, and individual article pages. That distinction keeps preservation evidence separate from the editorial decisions needed for a navigable archive.

Canonical records retain slugs, volume and edition identifiers, dates, categories, authors, excerpts, bodies, image relationships, original URLs, format, and sort keys. Stable slugs prevent later title edits from unexpectedly changing public URLs.

Illustrative, simplified code explaining the implementation pattern. This is not a copy of private production source.

// Illustrative canonical archive contract.
type ArchivedStory = {
  slug: string;
  title: string;
  volume: number;
  edition?: string;
  date?: string;
  category: string;
  authors: string[];
  originalUrl: string;
  images: string[];
  format: "article" | "slides";
  sortKey: string;
};

function validateArchive(stories: ArchivedStory[]) {
  const slugs = new Set<string>();
  for (const story of stories) {
    if (slugs.has(story.slug)) throw new Error("Duplicate canonical slug");
    slugs.add(story.slug);
    if (!story.authors.length) throw new Error("Missing attribution");
    if (!allowedCategories.has(story.category))
      throw new Error("Unmapped taxonomy value");
  }
}

Taxonomy and the challenge of mapping 300+ articles

The archive needs a controlled topic vocabulary rather than a new category for every variation in imported labels. Science, animals, art, culture, history, sports, and entertainment need consistent mapping while retaining the original source text and authorship.

Article identity and newsletter identity are different. Multiple contributions can belong to one volume, and a Google Slides edition can correspond to multiple canonical stories. The presentation catalog therefore stores associations to story slugs, rather than assuming a one-to-one relationship.

Unknown dates should remain unknown instead of receiving an invented timestamp. Imported material can have inconsistent formatting, repeated titles, partial images, and ambiguous authorship. Audit records and source-to-route manifests make those exceptions visible and allow targeted corrections without discarding the source archive.

Build-time routing and targeted interactivity

Build-time routes create permanent story, author, topic, and presentation pages. Search and filtering operate on prepared metadata in the browser, while the readable article is already present in HTML. React is loaded only on opted-in interactive pages.

Illustrative, simplified code explaining the implementation pattern. This is not a copy of private production source.

// Illustrative client-side archive filtering over canonical metadata.
function filterStories(stories, { query, category, author }) {
  const term = query.trim().toLocaleLowerCase();
  return stories.filter(story => {
    const searchable = `${story.title} ${story.excerpt ?? ""}`
      .toLocaleLowerCase();
    return (!term || searchable.includes(term)) &&
      (!category || story.category === category) &&
      (!author || story.authors.includes(author));
  });
}

Shared layouts generate canonical metadata, structured data, social cards, navigation, and the footer. Google Slides URLs are validated before constructing an iframe. Thumbnail associations and route manifests keep presentations discoverable without forcing a visitor to open an external deck first.

Media preservation, QA, and long-term maintainability

Media import and synchronization scripts maintain local images and thumbnails. Sharp is available for image processing, and audit files record migration and image checks. The archive build includes an SEO output check, while Playwright and axe-core support browser and accessibility validation.

Meaningful checks compare content counts, detect duplicate slugs, validate author and category references, inspect broken image paths, and verify sitemap coverage. The goal is preservation fidelity as well as visual quality: a successful build should not silently drop a child’s contribution.

Project impact
300+ articles

The body of work that motivated the preservation project, based on our family’s archive.

The result is a browsable, maintainable archive of my children’s writing, research, artwork, and newsletters. The value is preserving their work and making it accessible over time, rather than attaching a commercial revenue claim to a family project.