Website Crawling

If you already have a website, you can automatically crawl its content and import it into your assistant's knowledge base instead of copying it page by page. HaveAI extracts and processes the text from your site's pages.

How It Works

HaveAI uses your site's sitemap to discover your pages, then extracts the text content from them and adds it to your knowledge base. This is the most practical way to import content from multi-page sites quickly.

Starting a Crawl

1

Go to Web Crawling

Open AI Training β†’ Website Crawling in the panel.
2

Enter the site address

Enter the address of the site to crawl (e.g. https://yoursite.com).
3

Select pages

HaveAI lists the pages it found. Select the ones you want to import.
4

Import

The content of the selected pages is extracted, processed and added to your knowledge base.

Some sites may block access

Sites with protections like Cloudflare, a WAF (firewall) or password protection may block automatic crawling. In that case HaveAI shows you a descriptive error. You can then upload the content as a document or add it manually.

Review the crawled content

Automatic crawling is fast, but site content may include irrelevant text like menus and footers. Reviewing your knowledge base after import and removing unnecessary entries improves answer quality.

When to Use It

  • If you have a content-rich existing corporate site (ideal for a quick start).
  • If your product/service pages are already detailed on your site.
  • If you want to bring your blog or help-center content to the assistant.

Secret keys are automatically redacted

If scanned pages contain sensitive data such as API keys, passwords, or private keys (for example, a token left in a code comment), HaveAI automatically detects and masks thembefore the content is processed. Secrets are never stored in the database or exposed in assistant replies.

Related Topics

Was this page helpful?