The single highest-leverage move for any author or publisher is producing a validated ONIX 3.x product record with complete descriptive fields for every edition. That combination determines whether retailers accept a title cleanly or bounce it back for correction. Getting it right hinges on standards work from EDItEUR and guidance from BISG, plus tools like Alhora that catch problems before a feed ever leaves your desk.
TL;DR:
- Using ONIX 3.x with complete, validated descriptive fields is crucial for avoiding rejections at the retail level and ensuring smooth acceptance.
- Correctly assigning detailed subject codes like Thema or BISAC and keeping codelists current reduces the risk of feed failures and listing errors.
- Accurate metadata for titles, descriptions, keywords, and contributor names enhances discoverability, especially when aligning with platform-specific rules.
- Properly declared pricing, rights, and accessibility features prevent suppression or delays in market listing, particularly for international distribution.
- Maintaining a standardized, validated metadata workflow across all formats and channels ensures consistency and minimizes manual correction efforts.
Table of Contents
- Book Metadata Best Practices: The Essentials Checklist
- What Standards Should You Use: ONIX, Thema, BISAC, or ISBN?
- How Do Titles, Descriptions, and Keywords Affect Discoverability?
- Which Commercial and Supply Details Keep a Title From Being Suppressed?
- What Accessibility Metadata Do Ebooks and Audio Need?
- How Do You Validate Metadata and Avoid Feed Rejections?
- What Are the Most Common Metadata Mistakes, and How Do You Fix Them?
- How Do You Manage Metadata at Scale Across a Catalog?
- What Privacy and Legal Considerations Apply to Book Metadata?
- How Does Metadata Function Differently Across Sales Channels?
- The Uncomfortable Truth About Metadata Advice
- Sources
Book Metadata Best Practices: The Essentials Checklist
Before a title goes anywhere near a distributor, run through what actually needs to be filled in. Retailers and aggregators reject files for gaps that take five minutes to fix.
- Identifiers: a unique ISBN or GTIN for every edition and format, tagged with the correct product-form code.
- Title fields: title, subtitle, and series position entered separately, never combined into one marketing string.
- Descriptions: a short blurb and a longer description, both written for humans, not just search engines.
- Contributors: every author, editor, illustrator, and translator listed with a defined role.
- Subjects: 3 to 5 specific Thema or BISAC codes, never a single broad category.
- Audience and language: age range or reading level where relevant, plus the correct language code.
- Physical specs: extent (page count), trim size, and cover image meeting platform resolution requirements.
- Price and territory: currency, tax treatment, and which markets the title is cleared to sell in.
- Accessibility flags and supply details like publication date and stock status.
Booksellers specifically ask for carton quantities and age/interest ranges on children's titles, according to feedback compiled by the Australian Publishers Association. Skip those fields and your book may simply not appear in ordering systems that filter on them.
What Standards Should You Use: ONIX, Thema, BISAC, or ISBN?
ONIX 3.x is the maintained international standard for exchanging book product metadata, and EDItEUR has effectively retired ONIX 2.1 support across most major trading partners. If your export tool still generates 2.1 files, you're feeding a format most retailers no longer fully process. Run every file through XSD or strict XSD validation before it goes out the door, since schema errors are the most common reason feeds get flagged.
Subject classification is where a lot of publishers underinvest. Send Thema codes if you're selling internationally, BISAC if a trading partner's system still expects it, and label each scheme explicitly inside the ONIX record so there's no ambiguity about which taxonomy a code belongs to. Aim for 3 to 5 codes that describe the book precisely rather than one generic catch-all category.
Codelists change. A code that validated cleanly two years ago can become deprecated, and a feed built against an outdated list will fail silently at the retailer's end, not yours. Checking your codelist version against EDItEUR's current release before a major export run costs almost nothing and prevents an entire batch from bouncing.
Get this layer right and the payoff shows up downstream: fewer rejected feeds, faster distributor onboarding, and far less manual back-and-forth chasing corrections.
How Do Titles, Descriptions, and Keywords Affect Discoverability?
Field-level discipline is where most self-published and small-press titles lose visibility, not because the book is wrong for its market, but because the metadata misrepresents it. Amazon's own KDP metadata guidelines require the title field to match the cover exactly and explicitly prohibit promotional language in that field. A subtitle reading "The Gripping Bestseller You Can't Put Down" will get flagged, not celebrated.
- Keep titles and subtitles literal and factual; marketing language belongs in the description, not the title metadata.
- Write a short blurb (roughly 200 to 400 characters) for retail summary views, and a longer description (up to 4,000 characters on most platforms) that can use basic HTML formatting for readability.
- Build your keyword list around 2 to 3 word phrases real readers type into search bars, then echo that same language naturally inside your long description.
- Name series consistently across every book in that series, and tag contributor roles using EDItEUR's List 17 codes so "editor" and "illustrator" aren't confused with "author."
- Spell your name identically on every single record. A middle initial on book three and not on book one splits your catalog into two authors in some retailer systems.
Pro Tip: Search your own book's title plus your name on two or three retailers before launch. If a different edition, a pirated copy, or a name variant outranks your actual product page, your metadata has a consistency problem worth fixing now.
Which Commercial and Supply Details Keep a Title From Being Suppressed?
Pricing and rights metadata is where technically correct books quietly disappear from markets they're supposed to be sold in. Each format, ebook, paperback, hardcover, and audiobook, needs its own product record with its own ISBN and product-form code. Bundling formats under one record is a common cause of downstream confusion at the distributor level.
- Every price needs a currency code attached; a bare number without one gets rejected or misread by conversion systems.
- Publishers selling into the United Kingdom should include the correct price composite, VAT treatment, and currency, since a missing tax code can stall a listing entirely.
- Sales rights and territory restrictions must be declared accurately. Vague or blank territory fields often cause retailers to hide a title in markets it's actually cleared for, not just the ones it isn't.
- If you're exporting print stock internationally, include HS/commodity codes and country-of-manufacture data. Customs systems use these fields, and an omission can delay a shipment as easily as a bad address.
What Accessibility Metadata Do Ebooks and Audio Need?
Accessibility features only help readers if the metadata says they exist. ONIX supports this through ProductFormFeature entries, which declare specifics like a navigable table of contents, described images, and a properly sequenced reading order. An EPUB accessibility checklist built around WCAG and EPUB standards is the practical way to confirm the file itself qualifies before you ever touch the metadata.
Here's the part that trips up a lot of otherwise careful publishers: a genuinely accessible EPUB file with no accessibility flags in its ONIX record is functionally invisible to retailers filtering for accessible titles. The file can meet every technical standard and still never surface in an accessibility-focused search, because the platform has no metadata signal to act on. Declaring the feature isn't optional paperwork; it's the only way a retailer's system knows to show it.
How Do You Validate Metadata and Avoid Feed Rejections?
Validation needs to happen in layers, not as a single pass at the end. Schema-only checks catch structural errors, but business-rule validation, Schematron checks where your tooling supports them, catches the errors that actually cause rejections: wrong element order, missing conditional fields, and codes that are structurally valid but contextually wrong.
- Run XSD or strict XSD validation against the current ONIX schema version.
- Apply business-rule checks that mirror your aggregator's specific requirements, since automated Schematron-style checks catch conditional errors that schema validation alone misses, according to EDItEUR's implementation guidance.
- Cross-check subject codes, territory codes, and currency codes against your current codelist release.
- Confirm no required element is empty. The Library of Congress flags empty elements and incorrect element order as recurring causes of downstream rejection.
The most frequent culprits behind bounced feeds: obsolete ONIX 2.1 structure, missing or overly broad subject codes, currency mismatches, and empty required elements. Timing matters as much as accuracy. Complete product metadata 8 to 16 weeks before publication so retailers can process and index it ahead of launch day, and schedule deliberate update blocks whenever price, territory, or contributor details change rather than patching records ad hoc.
What Are the Most Common Metadata Mistakes, and How Do You Fix Them?
Most metadata problems fall into a small handful of repeat offenders, and nearly all of them have a same-day fix.
- Marketing copy in the title field: move every promotional phrase into the description, leaving the title and subtitle strictly literal.
- Overly broad subject codes: replace a single generic category with 3 to 5 specific Thema or BISAC codes that actually describe the book's content and audience.
- Currency or territory errors: rebuild price composites with explicit currency codes and confirm sales-rights territories are declared, not left blank.
- Missing accessibility metadata or weak cover images: add ProductFormFeature declarations and check cover files against each platform's resolution minimums before upload.
Pro Tip: Fix subject codes and title-field violations first. They're the two errors most likely to actively suppress a listing, and both take minutes to correct once you know the specific field a retailer flagged.
How Do You Manage Metadata at Scale Across a Catalog?
A typical production stack looks like this: a metadata database or PIM system feeds an ONIX exporter, which feeds an aggregator or distributor, who then passes validated records to individual retailers. Smaller catalogs sometimes skip the PIM layer and enter data directly into aggregator portals, which works fine until you're managing more than a handful of titles and start losing track of which record is current.
A workable validation pipeline runs schema checks, business-rule checks, and a codelist sync before every export, then stages updates in blocks rather than pushing one-off edits that create version drift across retailers. Maintaining one canonical metadata source that exports validated ONIX 3.x, even if you're also using retailer web forms, prevents the fragmentation that happens when three different people update three different spreadsheets.
Alhora fits into this workflow at the validation layer: its checks flag metadata and formatting issues, from missing accessibility features to inconsistent series data, without ever writing or altering your descriptive copy. When evaluating any metadata tool, check for:
- Native ONIX 3.x export and validation
- Thema and BISAC code support with current codelists
- Accessibility flag support (ProductFormFeature entries)
- Audit logs so you can trace what changed and when
What Privacy and Legal Considerations Apply to Book Metadata?
Book metadata itself is largely commercial and bibliographic, not personal data, but a few fields carry real legal weight worth taking seriously. Contributor records often include a real legal name, which counts as personal data under frameworks like the UK GDPR when it's tied to identifiable individuals through contracts, payment details, or contact information stored alongside the metadata. If your metadata database also holds contributor emails, bank details for royalty payments, or home addresses for contracts, that data needs the same handling care as any other personal record, restricted access, defined retention periods, and a clear basis for processing it.
Rights and territory metadata carries its own legal exposure. Declaring sales rights you don't actually hold, whether through error or optimistic guesswork, can put you in breach of contract with a publisher, agent, or co-author who holds rights you've claimed. Territory metadata that contradicts your actual licensing agreement isn't just a technical mismatch; it's a contractual problem that surfaces the moment a distributor cross-checks your rights declaration against their records.
Pen names and pseudonyms need careful handling too. ONIX supports separate fields for a legal contributor name versus a public-facing pseudonym, and mixing them up, or exposing a legal name a pseudonymous author specifically wants kept private, is both a metadata error and a potential privacy breach. If you manage titles for multiple authors, build a habit of confirming which name variant belongs in the public-facing fields before every export, not after a complaint arrives.
How Does Metadata Function Differently Across Sales Channels?
Metadata doesn't behave the same way on every platform, and treating every channel identically is a common source of quiet underperformance. Amazon's KDP enforces strict title-matching and keyword rules, and titles that violate them risk suppression rather than just poor ranking, per Amazon's own guidelines. Library and institutional distribution systems, by contrast, weight subject classification and standardized descriptions more heavily than marketing-style keywords, since librarians and acquisition systems search by controlled vocabulary rather than conversational phrases.
International retailers and aggregators lean harder on Thema classification because it's designed for cross-market interoperability in ways BISAC, built primarily for North American shelving and search conventions, isn't. A title correctly classified for a US retailer using only BISAC codes may show up miscategorized or not at all in a European storefront that expects Thema. Audiobook platforms add another layer entirely, often requiring narrator credits and runtime metadata that print and ebook records don't need at all.
The practical implication is that a single "universal" metadata record rarely serves every channel equally well. The core identifiers, title, and description can stay consistent, but subject codes, keyword emphasis, and certain supply fields often need channel-specific tuning. Publishers running multichannel distribution get the best results by treating their canonical metadata record as a source to adapt per channel, not a one-size file to blast everywhere unchanged. A distributor or aggregator that understands these channel differences will typically flag mismatches before they become rejected listings.
Choosing the right formatting and export tool matters here too. Alhora's ebook formatting software is built to generate clean EPUB exports across the retail platforms authors actually use, which is one less variable to manage when your metadata already has enough moving parts.

The Uncomfortable Truth About Metadata Advice
Most metadata advice treats copywriting and technical validation as separate problems, one for marketing, one for IT. That split is exactly why so much metadata quietly fails: a beautifully written description sitting inside an ONIX file that fails XSD validation never reaches a reader at all. The copy and the schema succeed or fail together.

The conventional advice oversells keywords and undersells codelists. Everyone talks about the perfect keyword phrase; almost nobody checks whether their Thema codelist is even current. A stale codelist can quietly break an entire feed while the keyword strategy inside it is flawless, and no amount of clever description writing fixes a validation failure.
If you take one thing from this guide, prioritize the boring layer first: validated ONIX structure, current codelists, and consistent identifiers across every edition. Get that foundation solid before obsessing over blurb phrasing. A perfectly written description attached to a broken feed reaches nobody. A plainly written one attached to a clean, validated record reaches every retailer on your distribution list.
— James
Sources
Start with EDItEUR's ONIX overview for schema and codelist downloads, BISG's metadata best practices resources for US-focused guidance, and BIC's Books Across Borders guidance for cross-border rules. Join EDItEUR's mailing lists or your national book industry group for update alerts.
- EDItEUR — Overview of ONIX for Books
- BISG — Metadata Best Practices Working Group
- Amazon KDP metadata guidelines
