Features/Content Genome
Private beta

Content Genome

Deterministically extract typed content blocks, structured data, and FAQs from every page into a stable typed representation.

For: Agencies, BusinessesCategory: understand

Private beta — available to agency partners

This capability is live and working for agencies in GeoScript's private beta program. Learn about agency access →

What Content Genome does

The Content Genome is GeoScript's structured representation of a client's website content. It takes the raw page snapshots produced by Website Intelligence and parses them into a stable, typed data structure — content blocks, schema markup, FAQs, meta descriptions, headings, and service area signals. This typed representation is what the Optimization Engine reasons over to produce recommendations. The Content Genome is the structured bridge between a raw website and actionable AI optimization insights.

How it works

Deterministic extraction

The Content Genome extractor is deterministic — no randomness, no wall-clock dependency, no AI-generated content inside the extraction process. The same page content produces the same genome every time. This is a design constraint, not an accident: determinism makes the system auditable and makes recommendations reproducible.

Typed content blocks

The extractor outputs a typed structure. Each extracted element has a defined type: LocalBusiness schema, FAQ block, service area reference, meta description, H1 tag, H2 tag, body content, and others. Typed blocks mean the Optimization Engine can apply specific rules — for example, detecting a missing LocalBusiness schema is only possible if the extraction layer reliably identifies when one is present or absent.

What gets extracted

The Content Genome captures: structured data (JSON-LD schema blocks), FAQ sections (Question/Answer pairs), meta descriptions, heading hierarchy, service area signals (geographic references in content), and body content blocks. The extraction is scoped to the public page content — it does not process content behind authentication or JavaScript-rendered sections.

ContentGenomeV2

The current implementation is ContentGenomeV2, which ships with the Slice D build. It processes every snapshot produced by Website Intelligence and makes the output available to the Optimization Engine in the same pipeline run.

What Content Genome does not do

Clear boundaries increase trust. These limitations are documented explicitly, not buried in footnotes.

  • The Content Genome does not generate content — it extracts and classifies what already exists on the site.

  • The Content Genome does not render JavaScript — extraction operates on the static HTML snapshot from Website Intelligence.

  • The genome is not a semantic understanding layer — it is a typed structural extraction. It identifies the presence or absence of structured elements, not the quality of the prose.

  • GeoScript does not use the Content Genome to train AI models — extraction is a read-only analysis step.

Related

Other features

GeoScript is in private beta

Agency partners get access to the full capability set — acquisition, monitoring, website intelligence, optimization, deployment, and reporting — for $299/month per active client.