AI Crawler Access Policy
How East Baoyu separates search crawling, user-initiated retrieval, model-training access and aggregate crawl-health monitoring.
Effective and last reviewed: 20 July 2026
This page records how East Baoyu New Energy currently manages automated access to the public website. It separates search visibility, user-initiated retrieval and model-training access. A crawler being allowed by robots.txt does not guarantee crawling, indexing, ranking or citation.
Current crawler policy
| Token or crawler | Policy | Purpose and boundary |
|---|---|---|
| Googlebot | Allowed | Public Google Search crawling. Public pages must still return 200, remain indexable and meet the applicable search requirements. |
| Bingbot | Allowed | Public Bing search crawling and discovery. |
| OAI-SearchBot | Allowed | Automatic crawling for eligibility in ChatGPT search results. This is separate from GPTBot model-training access. |
| ChatGPT-User | Allowed | User-initiated page retrieval. OpenAI states that this agent is not used for automatic web crawling and robots.txt rules may not apply to an individual user action. |
| GPTBot | Disallowed | Model-training crawling is blocked until East Baoyu records and publishes a revised management decision. This does not block OAI-SearchBot. |
| Other public crawlers | Allowed subject to the wildcard rules | Public content may be fetched unless a more specific group applies. The WordPress administration area remains disallowed. |
The machine-readable rules are the authoritative access instructions: robots.txt. The current public URL inventory is declared at sitemap_index.xml.
Official references
- OpenAI crawler documentation, including the independent OAI-SearchBot, GPTBot and ChatGPT-User purposes.
- OpenAI OAI-SearchBot published IP ranges.
- OpenAI GPTBot published IP ranges.
- Google Search technical requirements.
Aggregate crawl-health monitoring
The website keeps a rolling 90-day operational aggregate for claimed crawler user-agent classes and WordPress-handled error responses. The aggregate stores the day, a crawler class, a route class, an HTTP status, a query-free normalized path, a count and the last-seen time.
The monitor does not store IP addresses, query strings, referrers, cookies or complete user-agent values. A user-agent label can be spoofed and is therefore reported as a claimed crawler class, not verified vendor identity. Where a security investigation requires identity verification, the appropriate official published IP ranges and Hostinger access records should be checked independently.
Coverage is limited to requests handled by WordPress. Static edge/CDN requests and requests blocked before WordPress are not included. The report is an operational supplement, not a replacement for Hostinger raw access logs, Google Search Console Crawl Stats or Bing Webmaster Tools.
Review and contact
The policy is reviewed when a crawler purpose changes, when a new automated-use decision is approved or when the website infrastructure changes. Questions and legitimate crawl-access problems can be reported to info@eastbaoyu.com.
Change record
- 20 July 2026: first controlled version; OAI-SearchBot and search crawlers allowed, GPTBot model-training crawl disallowed, 90-day aggregate crawl-health boundary published.