Content Genome
Deterministically extract typed content blocks, structured data, and FAQs from every page into a stable typed representation.
Private beta — available to agency partners
This capability is live and working for agencies in GeoScript's private beta program. Learn about agency access →
What Content Genome does
The Content Genome is GeoScript's structured representation of a client's website content. It takes the raw page snapshots produced by Website Intelligence and parses them into a stable, typed data structure — content blocks, schema markup, FAQs, meta descriptions, headings, and service area signals. This typed representation is what the Optimization Engine reasons over to produce recommendations. The Content Genome is the structured bridge between a raw website and actionable AI optimization insights.
How it works
Deterministic extraction
The Content Genome extractor is deterministic — no randomness, no wall-clock dependency, no AI-generated content inside the extraction process. The same page content produces the same genome every time. This is a design constraint, not an accident: determinism makes the system auditable and makes recommendations reproducible.
Typed content blocks
The extractor outputs a typed structure. Each extracted element has a defined type: LocalBusiness schema, FAQ block, service area reference, meta description, H1 tag, H2 tag, body content, and others. Typed blocks mean the Optimization Engine can apply specific rules — for example, detecting a missing LocalBusiness schema is only possible if the extraction layer reliably identifies when one is present or absent.
What gets extracted
The Content Genome captures: structured data (JSON-LD schema blocks), FAQ sections (Question/Answer pairs), meta descriptions, heading hierarchy, service area signals (geographic references in content), and body content blocks. The extraction is scoped to the public page content — it does not process content behind authentication or JavaScript-rendered sections.
ContentGenomeV2
The current implementation is ContentGenomeV2, which ships with the Slice D build. It processes every snapshot produced by Website Intelligence and makes the output available to the Optimization Engine in the same pipeline run.
What Content Genome does not do
Clear boundaries increase trust. These limitations are documented explicitly, not buried in footnotes.
The Content Genome does not generate content — it extracts and classifies what already exists on the site.
The Content Genome does not render JavaScript — extraction operates on the static HTML snapshot from Website Intelligence.
The genome is not a semantic understanding layer — it is a typed structural extraction. It identifies the presence or absence of structured elements, not the quality of the prose.
GeoScript does not use the Content Genome to train AI models — extraction is a read-only analysis step.
Related
Other features
Website Intelligence
Private betaFetch, snapshot, and analyze a client's public website to build the structured evidence base GeoScript reasons over.
Optimization Engine
Private betaTurn Content Genome findings into typed, deterministic recommendations proposing exact DOM-level changes to improve AI understanding.
GeoScript is in private beta
Agency partners get access to the full capability set — acquisition, monitoring, website intelligence, optimization, deployment, and reporting — for $299/month per active client.