Benchmark Method
\nWe tested PostgreSQL 16 on Ubuntu 24.04 using...
\nPrompt
How do I structure my blog posts (markup, metadata, headings) so AI assistants and LLMs can easily discover, understand, and cite them?
Latest observation
To make blog posts easy for AI assistants and LLMs to discover, understand, and cite, structure them like high-quality content for humans plus machine-readable metadata and clear semantics.
1) Use semantic HTML
Prefer meaningful tags over generic divs.
articlefor the postheaderfor title/date/bylinemainfor primary contentsectionfor logical sectionsh1for the post titleh2/h3for nested headingsnavfor table of contents and site navigationfigure+figcaptionfor images/chartsblockquotefor quotescode,prefor code samplesul/olfor liststablewith proper headers for comparisons/data
Example:
<article>
<header>
<h1>How to Tune PostgreSQL for Read-Heavy Workloads</h1>
<p>By Jane Doe • Published 2026-07-24 • Updated 2026-07-24</p>
</header>
<nav aria-label="Table of contents">
<ol>
<li><a href="#why-it-matters">Why it matters</a></li>
<li><a href="#key-settings">Key settings</a></li>
<li><a href="#benchmark-results">Benchmark results</a></li>
</ol>
</nav>
<section id="why-it-matters">
<h2>Why it matters</h2>
<p>...</p>
</section>
</article>
2) Put the key facts near the top
LLMs and search tools often benefit from early, explicit context.
Include near the beginning:
- what the post is about
- who it is for
- the main conclusion or recommendation
- the date published/updated
- your name or organization
- version/applicability if relevant
A concise intro paragraph helps a lot.
3) Use clear heading hierarchy
Headings should form a logical outline.
Good:
h1: one per page, the titleh2: major sectionsh3: subsectionsh4+: only if truly needed
Avoid:
- skipping levels arbitrarily
- using headings for styling only
- long, vague headings like “More thoughts”
Make headings descriptive:
- “How to choose a VPN” instead of “Things to know”
- “Benchmark results on M2 MacBook Air” instead of “Performance”
4) Add metadata in <head>
This helps discovery, indexing, and citation.
Use:
<title>: specific, descriptive<meta name="description">: short summary- canonical URL
- Open Graph tags
- Twitter/X cards
- author, publish date, updated date if supported
- robots directives if needed
Example:
<title>How to Tune PostgreSQL for Read-Heavy Workloads</title>
<meta name="description" content="A practical guide to PostgreSQL settings, indexing, caching, and benchmarking for read-heavy applications.">
<link rel="canonical" href="https://example.com/blog/postgresql-read-heavy-tuning">
<meta property="og:title" content="How to Tune PostgreSQL for Read-Heavy Workloads">
<meta property="og:description" content="A practical guide to PostgreSQL settings, indexing, caching, and benchmarking for read-heavy applications.">
<meta property="og:url" content="https://example.com/blog/postgresql-read-heavy-tuning">
<meta property="article:published_time" content="2026-07-24T10:00:00Z">
<meta property="article:modified_time" content="2026-07-24T12:30:00Z">
<meta name="author" content="Jane Doe">
5) Add structured data (JSON-LD)
This is one of the most useful things you can do.
Use schema.org BlogPosting, Article, or NewsArticle depending on the content.
Example:
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "BlogPosting",
"headline": "How to Tune PostgreSQL for Read-Heavy Workloads",
"description": "A practical guide to PostgreSQL settings, indexing, caching, and benchmarking for read-heavy applications.",
"author": {
"@type": "Person",
"name": "Jane Doe"
},
"datePublished": "2026-07-24T10:00:00Z",
"dateModified": "2026-07-24T12:30:00Z",
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://example.com/blog/postgresql-read-heavy-tuning"
},
"publisher": {
"@type": "Organization",
"name": "Example Media",
"logo": {
"@type": "ImageObject",
"url": "https://example.com/logo.png"
}
}
}
</script>
If relevant, also consider:
FAQPagefor question/answer postsHowTofor procedural guidesRecipefor recipesProductfor product reviewsDatasetfor data posts
6) Make citations easy
If you want LLMs to cite your post accurately, make source boundaries obvious.
Helpful practices:
- give each section a stable anchor ID
- use short, quotable paragraphs
- include tables with labeled columns
- use bullets for distinct claims
- cite your own sources with links
- state definitions clearly
- include date/version context
Example section with anchor:
<section id="benchmark-method">
<h2>Benchmark Method</h2>
<p>We tested PostgreSQL 16 on Ubuntu 24.04 using...</p>
</section>
That lets assistants refer to:
https://example.com/blog/postgresql-read-heavy-tuning#benchmark-method
7) Keep content machine-friendly
LLMs handle content better when it is:
- unambiguous
- specific
- well organized
- not hidden behind tabs, accordions, or heavy JS
- not embedded only in images
- accessible without requiring login
Avoid:
- text inside images
- important content loaded only after interaction
- duplicated near-identical pages
- excessive boilerplate around the real content
- vague pronouns without referents
8) Use alt text and captions
For images:
- add descriptive
alt - use captions for context
- explain charts in the surrounding text
Example:
<figure>
<img src="cache-hit-rate.png" alt="Cache hit rate increased from 72% to 94% after enabling shared_buffers tuning">
<figcaption>Figure 2. Cache hit rate before and after tuning shared_buffers.</figcaption>
</figure>
9) Publish a clean, crawlable page
Assistants often rely on what search engines and crawlers can access.
Make sure:
- the page is indexable
- canonical URL is correct
- the content is server-rendered or otherwise crawlable
- robots.txt doesn’t block important pages
- no broken structured data
- pages load reasonably fast
- there’s one clear primary page for each article
10) Provide an RSS/feed or sitemap
These help discovery.
- XML sitemap for indexing
- RSS/Atom feed for updates
- separate feeds for categories/tags if useful
11) Write in a style that is easy to quote
If you want citations, include concise statements that can be lifted accurately.
Good example:
“For read-heavy workloads, increasing shared_buffers helps only up to the point where cache misses stop decreasing materially.”
This is easier to cite than a buried implication across three paragraphs.
12) Suggested checklist
For each post, ensure:
- one clear
h1 - descriptive
title - good meta description
- canonical URL
- publish/update dates
- author name
- JSON-LD structured data
- logical heading hierarchy
- stable section anchors
- accessible images with alt text
- crawlable, server-rendered content
- sitemap/feed inclusion
Minimal template
Here’s a simple starting point:
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Post Title</title>
<meta name="description" content="One-sentence summary of the post.">
<link rel="canonical" href="https://example.com/post">
<meta name="author" content="Your Name">
<meta property="article:published_time" content="2026-07-24T10:00:00Z">
<meta property="article:modified_time" content="2026-07-24T10:00:00Z">
<script type="application/ld+json">{ ... }</script>
</head>
<body>
<article>
<header>
<h1>Post Title</h1>
<p>By Your Name • Published July 24, 2026</p>
</header>
<main>
<section id="intro">
<h2>Introduction</h2>
<p>Summary of the post.</p>
</section>
</main>
</article>
</body>
</html>
If you want, I can also give you:
- a blog post SEO + AI discovery checklist,
- a full HTML + JSON-LD template, or
- recommendations for Markdown/MDX-based publishing workflows.