Source paper
Make your site readable by AI agentsAbstract
Vercel and Ora's study shows that AI agents must reach and read a page before they can ground an answer on it: in controlled probes, an answer that appeared only after JavaScript ran was missed by the client that reads raw pages, and a server refusal, an HTTP 403 response, defeated every client. That failure comes from how pages are delivered, not from anything specific to software companies, and the hotel industry bodies AHLA and HTNG reach the same conclusion, so we treat these basics as transferable to hotel websites. The study's figures on where agents go are a different matter: docs pages sat behind 47% of 844 grounded answers in a sample dominated by software companies, and that does not carry over to hotels without hotel data. We separate visibility, whether an agent can read a hotel's own stay facts, from bookability, whether it can complete a reservation, and adopt the first while watching the second.
1 · Why this research has value
Agents must reach and read before they can answer
The core finding is that agents must reach and read content before they can answer or act. The controlled probes compare two kinds of client. A fetch-only client downloads the raw HTML and cannot run JavaScript, so it misses anything a script adds to the page later. A JavaScript-capable client loads the page in a real browser and sees what a person sees. Two hard failure modes stand out.
- 0 of 5
- When the answer only appeared after the page's scripts ran, an agent reading the raw page never found it. An agent using a real browser found it every time.
- 0 of 10
- When site security blocked bots with an HTTP 403, in two test setups, neither kind of agent got the answer once, however capable.
- 394 of 844
- Answers traced to a page, 47%, came from docs pages. The sites were mostly software companies, so we do not assume this holds for hotels.
The results are clean but the sample is small: one test site and five topics per setup. Because the failure sits in page delivery, the mechanism is a candidate for transfer to other industries.
The agent runs show three further patterns. Agents reached the homepage in 69% of 1,033 runs. It was their first step in 92% of those runs, and in 59% of the same runs the next hop was a docs page. Docs pages were the most common source of grounded answers, behind 394 of 844. Discovery files were found mostly by following links rather than by guessing their location, in 86–97% of attributed fetches depending on file type.
The study states its own sample limits: six sites account for 691 of 1,033 runs, or 66.9%, and the run data is observational and run-weighted. Fetchability failures look transferable as engineering fact. SaaS path frequencies are not proven for hotel sites. That is where our take begins.
2 · Our take
Seven takes for hotel websites
Reach before reason.
If the answer is not in returned HTML, or the server returns 403, the agent cannot ground. We expect the hotel equivalents to be client-rendered booking shells and bot-management walls. Whether they are common on hotel sites is unmeasured; section 4b covers this gap.
First-party canon beats agent-only side doors.
Most answers came from a real page. In 1,023 runs the study could trace where the answer came from; 844 of those, or 82.5%, came from a page the agent had fetched. llms.txt, Markdown mirrors as described by Cloudflare and RFC 7763, and JSON-LD help discovery or structured consumption. They do not replace a real page holding the answer in fetchable HTML.
Homepage is the router.
It is usually the first hop, and the next hop needs clear links to pages that answer. A hotel homepage that links only to a booking widget gives an agent nowhere to go.
Same surfaces humans use.
Agents in the study read homepage and docs-like pages, not a parallel site. Make guest-facing facts fetchable rather than building an agent-only mirror.
Discovery files work when linked, and off-path surfaces rarely get fetched.
86–97% of attributed discovery-file fetches came through links: 86% for llms.txt, 93% for .well-known and 97% for openapi.json. Agents fetched sitemap.xml directly in 4% of runs. A hotel llms.txt only pays if linked from pages agents already hit; sitemaps stay for search crawlers.
Visibility is not bookability.
Fetch-and-ground is visibility: can an agent find and read first-party stay facts. Completing a reservation through a booking engine, MCP, OpenAPI or a checkout protocol is bookability. Hotels can ship visibility before bookability, and mixing the two produces the wrong roadmap and the wrong success metric. However, the same rails may be used for both, such as supplying live availability, rates, and inventory data before completing a booking.
Sample skew is itself a take.
Docs dominance and Markdown-heavy harnesses were measured on SaaS and developer sites. What may transfer is the mechanism, not “hotels will get 47% of answers from /docs”. A linked, fetchable facts hub is still worth having on a hotel site; whether it should carry a /docs label and information architecture is open.
.well-known and agentic commerce
Agents reached a .well-known path in 234 of 1,033 runs, or 22.7%. That is a discovery path, not a commerce-protocol evaluation. Separately, the Universal Commerce Protocol publishes a merchant profile at /.well-known/ucp, as set out in the UCP specification and Google's write-up, and Google has announced UCP for Lodging with onboarding “coming soon”. That sits on the bookability side of take 6: .well-known is how an agent learns how to act once readable facts already exist. Shipping a UCP profile is not the same as shipping fetchable stay facts.
3 · Supporting evidence
Evidence independent of the source
| Evidence | What it says | How it changes their finding | Strength |
|---|---|---|---|
| AHLA/HTNG Strategic Brief: Leveraging MCP, Hotel Content, and AI Tools to Maximize Direct Bookingsv1, 20 Jan 2026. Industry association guidance developed by its Global Technology 100 group. Guidance, not measurement. | The US hotel industry body tells hotels to open crawler access, keep rooms, amenities and offers “indexable and publicly accessible”, add schema.org markup, and separately confirm the booking engine “can handle AI-driven handoffs”. It warns that if bots cannot access the site, LLMs “may not know your property exists”. | Independent hospitality-side arrival at the same reach-first conclusion, and at the visibility/bookability split, with crawl access as one track and MCP with live availability, rates and inventory as another. Moves takes 1, 2 and 6 from analogy to corroborated guidance. | Signal |
| Cloudbeds: How AI recommends hotels810 prompts across ChatGPT, Perplexity and Gemini, six destinations, 145 properties. Vendor research with stated sample and method; commercial interest in the result. Citation share, not fetchability. | OTAs received 55.3% of all citations; official property websites 13.6%. 72.4% of recommended properties were brand-affiliated. | Shows the first-party visibility gap is live in lodging: agents are answering hotel questions from OTAs, not property sites. It does not say why, so it supports the problem, not Ora's mechanism. | Signal |
| Google Search Central: JavaScript SEO basicsDocumentation from the operator of the largest crawler. | “Server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript.” | Independent corroboration of the JavaScript-only failure mode from outside Ora's probes. | Signal |
| Cloudflare: Content Independence Day1 Jul 2025. First-party statement of platform policy by the CDN in front of a large share of the web. | Cloudflare changed its default “to block AI crawlers unless they pay creators for their content”. | Makes the 403 failure mode structurally likely for any site on default settings, hotels included. Does not measure hotel incidence. | Signal |
| Google Search Central: AI features and your websiteDocumentation from Google. | Google Search “doesn't use” llms.txt; doing so “will neither harm nor help your site's visibility”. Structured data “isn't required for generative AI search”. | Tempers the llms.txt recommendation to optional at most, and supports treating JSON-LD as additive rather than the sole carrier, as in take 2. | Signal |
| web.dev: Design agent-friendly site UXUpdated 1 Apr 2026. Documentation from the Chrome team. | Recommends semantic HTML, stable layouts, ARIA roles, and real button and anchor elements because “agents recognize these as interactive”. | Supports take 4: agents work best on the same well-built pages humans use, not on a separate surface. | Signal |
| schema.org HotelA subtype of LodgingBusiness. Community standard hosted at schema.org, verified directly. | A published vocabulary for exactly the stay facts guests ask about: checkinTime, checkoutTime, petsAllowed, amenityFeature, numberOfRooms. | Gives hotels a structured carrier for first-party facts, as in take 2, without inventing an agent-only format. | Confirmed |
| Google: UCP for LodgingGoogle developer documentation; announcement stage. | UCP “now unifies digital Hotel Booking”; “detailed onboarding and specs coming soon”. | Confirms bookability is a separate, still-forming track, as take 6 argues. Nothing a hotel can ship today. | Signal |
| Google: New ways to plan and book travel in AI Mode27 Aug 2026. First-party product announcement. Hotel booking rolling out in the US, in English. | Hotel booking inside AI Mode through named partners: Booking.com, Expedia, Hotels.com, Priceline, Trip.com, Hilton, Marriott, IHG, Choice and Wyndham. Payment runs through Google Pay, and “the hotel or booking platform will act as the merchant of record”. | Agent-completed hotel booking is live, but only for OTAs and large chains. Supports take 6: bookability is gated by partnership, while visibility is open to any hotel. The post does not say the flow runs on UCP. | Signal |
4 · Gaps and aspects to consider
What the evidence cannot yet carry
4a. In their study
- Run data is observational and run-weighted; six sites account for 66.9% of runs, so path frequencies describe those sites more than “the web”.
- The open dataset records the mix behind the 1,033 runs: four models in two harnesses. The models were claude-haiku-4-5 in 327 runs, claude-sonnet-4-6 in 236, gpt-5.4 in 236 and claude-fable-5 in 234; the harnesses were claude-agent-sdk in 541 runs and eve in 492. Three of the four models are Anthropic's and no consumer assistant was run as a product, so path and Markdown-request patterns may reflect these two harnesses more than agents in general.
- Answer provenance is attributed by the study's own method; the guide describes it but no third party has audited it.
- The source is a living knowledge-base page with a changelog. Numbers cited here are as of 8 Sep 2026.
- The four models have since been superseded. The probe failures, on script-only answers and on 403s, come from the fetch tool, not the model, so a newer model does not change them. Which pages agents visit, how many links they follow and which files they look for could change, and need rerunning on current models.
4b. In transferring it to the hotel space
- No 2026 primary study of agent browsing on hotel websites was found in the sources reviewed. Transfer of docs-path frequencies to hotels is argument by mechanism and analogy, not measured hotel path share.
- No AuraScope-owned probe data yet on hotel bot walls, client-rendered booking widgets, or rate-calendar fetchability. A follow-on AuraScope study, in progress, will close this.
- Hotel UCP adoption and agent hit rates on lodging .well-known/ucp profiles are unmeasured; onboarding is not yet open.
5 · Hospitality applicability
What it asks of a hotel
Requirement keywords follow RFC 2119. Each SHOULD or MUST below traces to the source column or to a row in section 3. Where the hospitality-side claim has no established source, the row says Hypothesis.
| Their finding, by take | Hospitality claim | Hospitality-side source | Strength |
|---|---|---|---|
| 1aJavaScript-only answers and 403 block retrieval | Hotel sites MUST serve stay facts, such as check-in and check-out, cancellation, parking, pets and accessibility, in initial HTML, and SHOULD allow reputable agent and crawler user agents through bot management. | AHLA/HTNG brief: open crawler access; content “indexable and publicly accessible”. Google JavaScript SEO basics: “not all bots can run JavaScript”. | Signal |
| 1bPrevalence on hotel sites | Client-rendered booking shells and bot walls are common on hotel sites. | None found. Cloudflare's default AI-crawler blocking makes bot walls plausible but is not hotel-specific. | Hypothesis |
| 2Grounded answers trace to fetched first-party pages | Hotels SHOULD keep stable URLs for policies, amenities, accessibility, check-in and check-out, parking, pets and contact, and MAY add schema.org Hotel markup as an additive carrier. | AHLA/HTNG checklist: structured data for rooms, rates, amenities, location. schema.org Hotel. | Signal |
| 3Homepage is the first hop for most runs | A homepage that links only to a booking widget gives agents nowhere to go; linking to the facts pages is the likely fix. | None hospitality-specific. web.dev agent UX supports real anchor elements in general. | Hypothesis |
| 4Agents read the same pages people do | Guest-facing pages SHOULD use semantic HTML and real links and buttons. Hotels SHOULD NOT build an agent-only mirror as a substitute for a readable site. | web.dev agent UX, Chrome team. | Signal |
| 5Discovery files work when linked; sitemap fetched directly in 4% of runs | A hotel MAY publish llms.txt or Markdown mirrors, and if it does they MUST be linked from pages agents already hit. Sitemaps stay for search crawlers. | Google AI features guide: Search does not use llms.txt. Cloudflare Markdown for Agents. | Signal |
| 6Fetch-and-ground is separate from “enable agents to act” | Hotels SHOULD make their facts readable before enabling agentic checkout: visibility before bookability. Booking-engine handoff, MCP and UCP are a separate programme. | AHLA/HTNG brief: dual strategy of crawler access plus real-time data via MCP; “confirm your direct booking engine can handle AI-driven handoffs”. Google UCP for Lodging: onboarding coming soon. | Signal |
| 7aDocs pages behind 47% of grounded answers in a SaaS sample | Hotel information architecture rarely resembles a developer docs tree; a linked facts hub may use a /docs-style path, but SaaS docs frequencies should not be assumed to transfer. | Cloudbeds study: property sites received 13.6% of AI citations versus OTAs 55.3%, which shows the first-party gap but not its cause. | Hypothesis |
| 7bMarkdown requested frequently by study harnesses | Hotel CMSs rarely negotiate text/markdown; HTML-first, mirrors optional later. | None found for hotel CMS behaviour. | Hypothesis |
6 · Position
Adopt retrieval hygiene, watch the docs-first playbook
Adopt retrieval hygiene now:
- Serve stay facts in the initial HTML, server-rendered or prerendered.
- Keep those pages reachable, returning honest status codes.
- Make the homepage route to those pages.
- Publish discovery files only when they are linked from pages agents already visit.
- Score visibility separately from bookability.
These recommendations rest on Signal-grade evidence from three independent directions: Ora's probes, Google's crawler documentation and AHLA and HTNG's hospitality guidance. That is why they carry SHOULD and MUST in section 5.
Watch the SaaS docs-first playbook: a /docs-style hub, Markdown negotiation and OpenAPI or MCP as the default act path are one candidate shape for how agents understand a site, measured where the measurements were convenient. We do not assume SaaS path frequencies describe hotel sites.
7 · Open questions
What we do not know yet
- Is there an emerging standard for how agents understand a website, and is it SaaS-shaped, lodging-native, or not yet formed? This paper keeps that open on purpose.
- What share of hotel sites serve stay facts only after JavaScript, and what share challenge or 403 non-browser clients? The follow-on study addresses this.
- Does a linked facts hub, /docs-labelled or not, raise retrieval versus homepage-only and booking-widget paths on hotel sites? The follow-on study addresses this.
- Do the consumer assistants guests actually use, meaning ChatGPT, Gemini and Perplexity as products rather than models inside a research harness, traverse hotel sites the way the study's two harnesses did?
- How many hotels publish a .well-known/ucp profile, and do agents fetch it? Blocked until UCP for Lodging onboarding opens.
8 · References
20 sources
- Vercel with Ora.ai research lab. Make your site readable by AI agents. 28 Aug 2026. Accessed 8 Sep 2026.vercel.com/kb/guide/make-your-site-readable-by-ai-agents
- AHLA and HTNG. Strategic Brief: Leveraging MCP, Hotel Content, and AI Tools to Maximize Direct Bookings, v1. 20 Jan 2026. Accessed 9 Sep 2026.ahla.com/sites/default/files/MCP%26DB_AI-Strategic-Brief.pdf
- Cloudbeds. How AI recommends hotels, research data. Accessed 9 Sep 2026. Publication date not shown on page.cloudbeds.com/hotel-ai-recommendations/data/
- Google Search Central. Understand JavaScript SEO basics. Accessed 9 Sep 2026.developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics
- Google Search Central. AI features and your website. Accessed 9 Sep 2026.developers.google.com/search/docs/fundamentals/ai-optimization-guide
- Google Developers. UCP for Lodging. Accessed 9 Sep 2026.developers.google.com/hotels/ucp
- Google, The Keyword. New ways to plan and book travel in AI Mode. 27 Aug 2026. Accessed 30 Sep 2026.blog.google/products-and-platforms/products/search/book-travel-ai-mode/
- Google Developers Blog. Under the hood: Universal Commerce Protocol. 11 Jan 2026. Accessed 8 Sep 2026.developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp/
- UCP. Specification overview, 2026-04-08. Accessed 8 Sep 2026.ucp.dev/2026-04-08/specification/overview/
- Cloudflare. Content Independence Day: no AI crawl without compensation. 1 Jul 2025. Accessed 9 Sep 2026.blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/
- Cloudflare Docs. Markdown for Agents. Accessed 8 Sep 2026.developers.cloudflare.com/fundamentals/reference/markdown-for-agents/
- Chrome team, web.dev. Design agent-friendly site UX. Updated 1 Apr 2026. Accessed 9 Sep 2026.web.dev/articles/ai-agent-site-ux
- schema.org. Hotel. Accessed 9 Sep 2026.schema.org/Hotel
- W3C. JSON-LD 1.1.w3.org/TR/json-ld11/
- IETF. RFC 7763: The text/markdown Media Type.rfc-editor.org/info/rfc7763/
- llmstxt.org. The /llms.txt file.llmstxt.org/
- AgentReady, Ora.ai research lab. Agent runs dataset README, CC BY 4.0. Accessed 9 Sep 2026.github.com/agentready-org/standard/blob/main/data/README.md
- IETF. RFC 2119: Key words for use in RFCs to Indicate Requirement Levels.rfc-editor.org/info/rfc2119/
- IETF. RFC 8615: Well-Known Uniform Resource Identifiers.rfc-editor.org/info/rfc8615/
- OpenAPI Initiative. OpenAPI Specification. Accessed 25 Sep 2026.spec.openapis.org/oas/latest.html
Key terms
In plain language
- AI agent
- Software, such as ChatGPT, Gemini or Perplexity, that reads websites and acts on a person's behalf.
- Grounding
- Basing an answer on a page the agent actually read, rather than on what the model already remembers.
- Answer provenance
- Where an answer came from. The source study traced each answer back to the page it was based on, where it could.
- Harness
- The software that wraps a model and gives it tools such as fetching web pages. The source study used two.
- Fetch-only client
- An agent tool that downloads a page's raw HTML without running its JavaScript. Anything the page adds later with JavaScript is invisible to it.
- JavaScript-capable client
- An agent tool that runs a real browser, so it sees the page as a person does.
- Initial HTML, server-rendered
- The page content the server sends before any JavaScript runs. Facts in the initial HTML can be read by every client.
- HTTP 403
- The status code a server returns when it refuses to serve a page, often because bot-protection software has flagged the visitor as a bot.
- Bot management
- Security settings, usually part of a firewall or content delivery network such as Cloudflare, that decide which automated visitors are let through.
- Stay facts
- The details guests ask about before booking: check-in and check-out times, cancellation, parking, pets and accessibility.
- SaaS
- Software sold as an online service. Most sites in the source study were SaaS or developer-tool companies.
- Discovery files
- Files such as llms.txt, sitemap.xml or openapi.json that tell machines what a site contains.
- llms.txt
- A proposed text file that lists a site's key pages for AI models.
- OpenAPI
- A standard format for describing a web API so software can call it, originally based on the Swagger Specification. A site's openapi.json is its description in this format.
- Markdown mirror
- A plain-text copy of a page, served to agents alongside the normal page.
- Structured data
- Labels in a page's code that tell machines what each fact means, written with the schema.org vocabulary in the JSON-LD format.
- .well-known
- A standard location on a website, defined in RFC 8615, where machines look for configuration files.
- MCP
- Model Context Protocol: a standard way for an agent to call a business's systems directly, for example to check availability.
- UCP
- Universal Commerce Protocol: a standard for agents to complete purchases. Google has announced a version for lodging.
- OTA
- Online travel agency, such as Booking.com or Expedia.
- MUST, SHOULD, MAY
- In capitals these follow the IETF convention in RFC 2119: MUST is a requirement, SHOULD a strong recommendation that allows good reasons to differ, MAY is optional. They say what we recommend, not how sure we are.
- Hypothesis, Signal, Confirmed
- How strong the evidence is. Hypothesis: plausible but without a supporting source. Signal: indicative evidence from a credible source. Confirmed: reproduced by AuraScope or agreed by two independent sources.
