Can Claude read your website at all? The first check for Malaysian SMEs
Co-Founder, Geo One

Meta description: Before spending money on AI visibility, check whether Anthropic's crawlers can fetch your site at all. A ten-minute robots.txt, server-header and JavaScript check for Malaysian SMEs.
Can Claude read your website at all? The first check for Malaysian SMEs
To appear in Claude answers in Malaysia, three things need to be true at the same time. Anthropic's crawlers have to be able to reach your pages. Your business facts have to match across your website, your Google Business Profile and the directories you are listed in. And other sources have to describe you in roughly the same way you describe yourself. Miss any one of those and Claude may still answer the question, using someone else's information. Nobody can guarantee an assistant names a particular business, but you can fix the signals it reads.
They are not equally urgent. The reachability check comes first, because if Claude cannot fetch your page at the moment someone asks, every hour spent on wording, schema and directory listings is spent on a page the assistant never sees. This article covers that first check in full and points you to the other two at the end.
Written for owners of Kuala Lumpur professional services firms with 2 to 50 staff: clinics, law and accounting practices, property agencies, freight forwarders. It comes from Geo One, a Generative Engine Optimisation (GEO) and Answer Engine Optimisation (AEO) agency in Kuala Lumpur. If the terms are new, start with our explainer on what answer engine optimisation is.
Where Claude's answers come from
Claude is the AI assistant built by Anthropic. Malaysian buyers use it the way they used to use a search box. "Find me a freight forwarder in Port Klang who handles LCL shipments." "Which KL accounting firm does SST filing for a small trading company?" The answer arrives as a short paragraph with a few names in it, not ten blue links.
Claude draws on two sources. One is what the model learned during training, which you cannot edit. The other is what it fetches from the live web at the moment somebody asks, which you can influence. That live retrieval path is where GEO work applies, and it is the only part of the system a small firm can audit for itself.
Anthropic does not use one crawler
This is the part most owners have never looked at, and the reason well-built sites go missing.
Anthropic runs separate user agents for separate jobs. At the time of writing its published list includes ClaudeBot, associated with training data collection; Claude-SearchBot, associated with search indexing; and Claude-User, which fetches a page in real time when somebody asks Claude to look at something. The older names anthropic-ai and Claude-Web still appear in plenty of robots.txt files written in 2023 and 2024.
Allowing one and blocking the others is the common failure. A site that permits the training crawler but blocks the search-time or user-initiated agent looks cooperative in a robots.txt audit and is still invisible at the exact moment a buyer asks a question.
The reverse also happens. An owner who read about AI scraping blocks everything Anthropic-shaped, then wonders why the firm is never mentioned. Blocking training and allowing retrieval is a defensible position; blocking both by accident is not.
Anthropic maintains its own page on how its crawlers behave and how site owners can control them. Check the current names there before you edit anything, the list has changed before and will change again.
The ten-minute check
1. Read your robots.txt. Open yourdomain.com/robots.txt in a browser. You are looking for two things: a blanket User-agent: * followed by Disallow: /, and any rule naming the agents above. If your developer added AI-blocking rules after the scraping coverage a couple of years ago, a lot did, quietly, without asking, they are probably still sitting there.
2. Check what your server actually returns. robots.txt is a request, not a wall. Your host or CDN may refuse the connection regardless of what the file says. Ask whoever manages your hosting to fetch a key page as each Anthropic user agent and report the result:
curl -A "ClaudeBot" -I https://yourdomain.com/services
curl -A "Claude-User" -I https://yourdomain.com/services
A 200 is what you want. A 403 or 429 means something is blocking or throttling at the edge. An X-Robots-Tag: noindex header means the page loads but is marked not to be used.
Two culprits show up repeatedly on Malaysian SME sites. The first is Cloudflare's AI-crawler blocking: Cloudflare introduced a one-click block for AI bots in July 2024 and began blocking AI crawlers by default for new domains from July 2025, so a site moved onto Cloudflare recently may be blocking without anyone having chosen to. The second is security plugins and WAF rules on shared hosting, which often treat an unfamiliar user agent or a data-centre IP range as an attack and return a 403 with no log entry anyone ever reads.
3. Look at the page with JavaScript off. View source on your services page and search for your phone number, your suburb and the name of the service you want to be found for. If those words only appear once scripts have run, assume a retrieval agent may not see them. Sites built on page-builder templates, where the service descriptions sit inside tabs or accordions, are the usual offenders.
If all three pass, Claude can read you. That is the floor, not the finish line.
If you find a block, what to change
You do not need a new website. In most cases it is three edits.
In robots.txt, remove the outdated agent names and make the permission explicit. Robots.txt defaults to allow, so you only need named rules where a blanket Disallow would otherwise catch them:
User-agent: ClaudeBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
Keep any genuine exclusions, admin paths, client portals, staging subdomains, internal PDFs with client names in them. Reachability is not the same as publishing everything.
At the edge, turn off the blanket AI-bot rule in Cloudflare or your WAF and re-run the two curl commands. If your host cannot tell you what rule fired, that is worth knowing about your host.
On the page itself, move the facts that matter, services, suburbs, phone number, opening hours, the SSM-registered entity name, into plain server-rendered text. Then re-check. The whole job is usually under an hour of developer time, and it is worth insisting on a before-and-after status code rather than a "done" in WhatsApp.
Then the other two
Only once the page is reachable do the remaining checks pay for themselves.
Facts that match everywhere. Claude weighs corroboration. If your website says Jalan Sultan Ismail, your Google Business Profile says Jalan Raja Chulan and a 2019 directory listing carries a dead landline, the assistant has three versions of you and no reason to trust any of them. The details that quietly break this in Malaysia are ordinary ones: the SSM-registered name against the trading name, a unit number written four ways, "Jalan" against "Jln", a mobile number given as 012-345 6789 in one place and +60 12-345 6789 in another. Pick one form of each and push it everywhere. Our six-check Google Business Profile audit method walks through the order to do it in.
Descriptions written by other people. An assistant is more likely to name a firm that third parties describe in the same terms the firm uses about itself, the same specialisations, the same locations, the same client types. That is a citation and mention building problem rather than a website problem, and it is the slowest of the three to move. If Claude already describes you but describes you wrongly, that is a different repair job again: see what to do when an AI assistant gets your business wrong.
Start with step one this week. It costs nothing, it takes under ten minutes, and it is the only one of the three that can silently cancel out all the others.
Frequently asked questions
How do I get my business mentioned in Claude's answers?
Three things need to be true at the same time: Anthropic's crawlers can reach your pages, your business facts match across your website, Google Business Profile and directories, and other sources describe you roughly the way you describe yourself. Miss one and Claude may still answer the question using someone else's information. Nobody can make an assistant name a particular business, but you can fix the signals it reads.
Where does Claude actually get its answers from?
Claude draws on two different things. One is what the model learned during training, which you cannot edit. The other is what it fetches from the live web at the moment somebody asks, which you can influence. That live retrieval path is where GEO work actually applies. Anthropic sets out how the platform and its crawlers behave in its own documentation.
My site is not blocked in robots.txt, so Claude can read it, right?
Not necessarily. Anthropic does not use a single crawler. There is a crawler associated with training, a separate one used for search, and a third that fetches a page when a user asks Claude to look at it. Allowing one while blocking the others is a common mistake. A site that permits the training crawler but blocks the search-time agent can still be unreadable at the exact moment a buyer asks a question.
What should I check first before anything else?
Check whether Claude can read your website at all. This is the first thing to look at, and it is the one most owners have never checked. Because Anthropic runs separate crawlers for training, for search and for fetching a page on request, you need all of the relevant ones allowed, not just one.
Is this relevant for a small firm in Kuala Lumpur?
Yes. This is written for owners of Kuala Lumpur professional services firms with 2 to 50 staff: clinics, law and accounting practices, property agencies and freight forwarders. Malaysian buyers now ask Claude things like which KL accounting firm handles SST filing for a small trading company, and the answer comes back as a short paragraph with a few names in it rather than ten blue links.
Can you guarantee Claude will recommend my company?
No. Nobody can make an assistant name a particular business. What you can do is fix the signals it reads: crawler access to your pages, consistent business facts across your website, Google Business Profile and directory listings, and third party sources that describe you in roughly the same way you describe yourself.

Bernard Leong
Co-Founder, Geo One
Nearly 20 years across energy, capital strategy and applied AI, including large-scale operational data at BP. Founded SkillsMe and Cryptrain.
More about the teamCheck Your AI Visibility
See how AI platforms currently view your business with a free scan.
Free AI Scan