How a crawl starts#
Adding a website through Websites → Start automatic setup automatically queues its first crawl from https://your-domain/. Later, an owner or admin can open that website's Knowledge tab and choose Crawl again.
The crawl is a background job. A queued status is not a completed import, and no fixed completion time is guaranteed. The website view refreshes processing status while visible; its refresh control lets you request the latest state manually.
What can be collected#
The normal crawler reads public HTML, extracts the main readable text where possible and follows links on the same hostname. It removes page scripts, styles and common navigation or footer elements from the answer source. It skips PDF, archive, image and media file links rather than extracting those files as documents.
Normal collection examines up to 60 pages per run. If normal collection produces no usable pages, a rendered-browser fallback can examine up to 30 pages. These are crawl limits, not a guarantee that the same number of pages will be reached or processed successfully.
Public access and hostname boundaries#
Reesponder does not log into protected customer accounts or administration areas to read their contents. A login screen, website password, anti-bot challenge or inaccessible page can prevent collection. Same-host redirects are followed within a bounded redirect limit; a redirect to another hostname can be rejected.
For normal HTML collection, the crawler checks wildcard disallow paths from robots.txt when available. It blocks private or local network targets and limits response sizes. Do not assume the crawl is a complete audit of every page or every access rule: verify the active source documents in the portal.
Read processing states#
In the website's Knowledge tab, the crawl shows its state, processed-page count and a visible error when available. In the Knowledge library, the website source moves through pending and processing to ready, or failed.
Pages processed counts pages with usable title and content. Knowledge documents identifies active stored documents, including other entries where relevant. A zero count needs investigation even if the website itself looks fine in your own signed-in browser.
Check coverage and make gaps explicit#
Open the website crawl source and use Open original page to compare collected text with important customer pages. Check policies, product details, service areas and contact information. Pages reachable only through search, infinite scroll or authentication may not appear in the crawl.
If a page is missing, publish accessible, linked HTML content or add a concise verified manual entry. If your website changes, use Crawl again; there is no customer-configurable automatic re-crawl schedule. For stock, account status or other live data, configure an Action rather than relying on a page snapshot.
Worked example: a footer-only contact detail is missing#
A business's phone number appears only in the repeated footer of a redesigned website. The site looks complete in a browser, but normal extraction removes common footer material from its answer source. The reviewer opens the stored documents and sees that the main customer-service text does not contain the number.
The maintainer publishes an accessible contact page with the approved details and links to it from the public site, or creates a verified manual fact in the correct scope. They request the supported crawl from the website's Knowledge tab and check the stored contact text afterwards. There is no arbitrary start-page selector in this flow: the connected website's crawl begins from its configured public root and follows eligible same-host links.
Record a coverage finding at document level#
Record each essential topic, the expected public URL and whether its relevant text appears in an active document. A missing image or layout element is not a failure to collect business text, while an omitted policy condition can be consequential even if many pages were processed.
For a gap, note whether it is an access, link, extraction or format issue and choose the appropriate remedy. Keep private pages protected. When a manual fact is the deliberate alternative, record that decision so the next maintainer does not repeatedly crawl the same unsupported attachment. The snapshot comparison below catches regressions that a total page count would hide.
Make a coverage checklist before judging the run#
List the pages essential to useful answers: contact information, service area, delivery, returns and representative products or services. Keep the list focused enough for somebody to compare it with the collected source reader. A large raw count does not prove these pages were included. A small source can still be useful if it covers the questions your business intends to support in its first release.
Check the public access path for each essential page. Is it reachable through ordinary same-host links, or only through search, login or an attachment? A policy available only as a PDF is not an eligible HTML document. If appropriate, publish its customer-facing text as accessible HTML and review the result. Otherwise create a verified manual fact containing only information suitable for customer answers. Do not expose private records to make collection easier.
Run the crawl through the connected website workflow and compare the result with your checklist. The portal does not offer an unlimited arbitrary bulk importer. A page outside the collection bounds can be missing without the whole job failing. Record the missing topic and chosen remedy. This creates a manageable content plan rather than a vague request to import everything on a large website.
Why your browser and stored text can disagree#
Your signed-in browser can show content unavailable to a public collector. A storefront password, administrator session or anti-bot challenge can change the returned page. A visual page may rely on scripts to insert information later. Normal collection reads eligible public HTML first. Rendered fallback is used when normal collection yields no usable content, not as a guaranteed second pass over every partially collected page.
Compare the stored document with the original page in a visitor session. Look for the main factual text, not simply whether the design looks complete. Information inside an image is not readable page text. Common navigation and footer elements are removed during extraction, so an important fact placed only in a repeated decorative footer may need a proper public policy page or manual entry. Pixel similarity does not establish text coverage.
Redirects can explain another mismatch. A configured host redirecting to a different hostname can be rejected although your browser follows it casually. Confirm the final public host and website record. Private-network protections also reject internal crawl targets. An access error should lead to a domain or website-access investigation, not indiscriminate removal of security controls. Keep public business information accessible while preserving protection around genuinely private systems.
Compare the active snapshot after refresh#
After a changed successful crawl, inspect the active documents again. The new snapshot can have different coverage from the old one. If a redesign removed links to a policy, a crawl may collect useful content but omit that policy. Success does not merge every formerly available page indefinitely into the new version. Keep the essential-page checklist so the difference is visible.
A wholly failed run can leave earlier usable documents available. That preserves previous content but does not make it fresh. Note the failed run and the age of the actual information. Do not report a policy update as complete simply because the old document still appears. Resolve public access or collection issues, inspect the subsequent active snapshot and test the changed answer. No fixed crawl completion time is guaranteed by a queued state.
For an urgent fact, a focused manual entry can support customer-facing answers while you investigate coverage. Review any conflicting source and choose scope deliberately. Avoid repeatedly creating near-identical copies of the policy; they increase maintenance without solving the cause. Finish with both the representative-answer test and missing-page comparison. Source state, processed-page count and successful widget installation each describe only part of the release.
Need help with this?
Tell us which website, guide and step you are working on. Keep passwords and private customer details out of the message.