I like opening a company website I have never seen before and trying to understand the business without touching anything else.

No sales call.

No internal documents.

No founder explaining what the product is supposed to do.

Just the website.

It is a surprisingly good exercise.

Within a few minutes you can usually work out what the company sells, who they think the customer is, what they consider important, what they are nervous about explaining, how complicated their pricing is and which questions probably appear in support every week.

Sometimes you can also see where the trouble is going to start.

A pricing table has a tiny footnote.

One product page calls something “included” while another calls it “available”.

An FAQ contains a very specific question that clearly did not get added there by accident.

There is a sentence like: “Existing customers will not be affected.”

Whenever I see something like that, I immediately want to know what happened.

Why did somebody decide that sentence needed to exist?

How many people asked before it was added?

What were they worried about?

A website is full of little fossils from previous customer questions.

That is one reason a public website is such an interesting starting point for a support system.

The website already contains years of decisions.

People sometimes talk about a website as if it were mainly marketing material.

That description feels too small to me.

A decent business website contains decisions from almost every part of the company.

Pricing tells you how the product is packaged.

Feature pages tell you what the company believes is worth selling.

Documentation explains how things actually work.

Terms and policies describe limits.

An implementation page shows what customers need to do after buying.

An FAQ often reveals the questions people kept asking after all the other pages were already published.

Even tiny wording choices tell you something.

If a company says “workspace” everywhere, that matters.

If they call customers “members” instead of “users”, that matters too.

If they consistently say “cancel anytime” rather than “no long-term contract”, you have learned something about the language they prefer.

A public website is a compressed version of the business. It leaves things out, sometimes important things, but there is far more knowledge in it than the homepage makes obvious.

This makes the first stage of teaching a support system unusually convenient.

The business has already written a lot of the material.

It probably spent months or years doing it.

Nobody called it a knowledge base while they were writing the product page, but parts of it function exactly like one.

A homepage tells me more than a list of features.

Imagine I open the homepage of a small software company.

The headline says: “Send proposals, collect signatures and get paid from one link.”

Already I know quite a lot.

This probably serves businesses rather than casual consumers.

Payments are part of the workflow.

Electronic signatures matter.

There is probably some concept of clients or recipients.

The company cares about reducing the number of tools involved in a job.

Then I scroll.

There is a section explaining that recipients do not need an account.

Interesting.

That is almost certainly a recurring customer question.

Another section says payments can be collected through Stripe.

Now I expect questions about Stripe accounts, fees, payouts and whether the company handles money directly.

The pricing page limits the cheapest plan to five active proposals.

Now the word active becomes important.

Does a completed proposal still count?

Does deleting one free a slot?

What happens when you reach six?

Nobody needed to give me a list of likely support questions.

The site created the list itself.

The boring pages are often incredibly useful.

I think people naturally pay attention to the flashy parts of a site first.

Homepage.

Product.

Pricing.

Those are useful.

The boring pages are where things get really interesting.

Shipping policies.

Returns.

Cancellation terms.

Technical documentation.

Account deletion instructions.

Privacy information.

Installation guides.

These pages contain details that customers care about once the simple marketing questions are over.

Suppose a clothing shop says returns are accepted within thirty days.

Fine.

Then its returns page says items must be unworn, with original tags attached, and final-sale items cannot be returned.

Suddenly a support system can answer a much more useful question:

“I tried the jacket on at home and removed it from the shipping bag, but the tag is still attached. Can I send it back?”

That answer comes from policy detail.

A homepage would never contain it.

This is why “learn from the website” sounds simpler than it actually is.

A website is rarely one source.

It is a collection of sources with different jobs.

URLs have a strange amount of meaning.

Even the structure of the site can help.

A page under: /docs/integrations/stripe probably has a different role from: /blog/why-we-love-stripe.

Both might mention Stripe twenty times.

I would trust them differently when answering a technical support question.

The documentation page is likely intended to describe current behaviour.

The blog article may be three years old.

It may describe a product version that no longer exists.

It may be educational rather than authoritative.

This becomes important very quickly.

A support system cannot treat every sentence it finds as equally valuable.

Humans do not do that either.

If I want to know whether a company still offers refunds, I trust the current refund policy more than a 2023 launch announcement.

We make these little judgments automatically.

Software has to make them deliberately.

Repetition is useful evidence.

Another thing I notice when looking through websites is repetition.

Companies repeat important information constantly.

“No credit card required.”

It appears in the hero.

Then pricing.

Then the FAQ.

Then underneath the signup button.

That repetition tells us something.

The company considers this information important enough to remove uncertainty before it appears.

If the same policy appears consistently across several current pages, confidence increases.

Repetition can also expose conflict.

Maybe pricing says thirty days of history while documentation says ninety.

Now we have a problem worth noticing.

A weak system picks whichever sentence happens to look relevant.

A careful system needs to notice that the business itself appears to disagree with itself.

Contradictions are valuable information. They tell you where blindly producing an answer would be risky, and they often reveal outdated pages the business did not realise were still public.

Dates become important sooner than you expect.

The web has a long memory.

Businesses do not always.

A company launches a feature in 2024.

It writes three articles about it.

In 2025 the feature changes completely.

The documentation gets updated.

The old blog posts remain.

In 2026 somebody asks how the feature works.

Search engines still happily return all four pages.

This is normal.

It also means a support system needs some sense of time.

Publication date can help.

Last-modified date can help.

Page type can help.

Explicit instructions from the business help even more.

I would rather have a system hesitate for a second than confidently explain a version of the product that disappeared two years ago.

Language is part of the knowledge.

There is another thing a website teaches that is easy to overlook: how the business speaks.

Imagine two companies selling essentially the same service.

One says: “Terminate your subscription.”

The other says: “Cancel whenever you like.”

Same underlying action.

Very different voice.

A support response that suddenly starts sounding like a lawyer or a corporate manual can feel foreign even when the information is correct.

People notice this more than they realise.

If a friendly little business suddenly replies: “Your request has been acknowledged and will be processed pursuant to our standard cancellation procedure,” something feels wrong.

The website provides examples of the company’s vocabulary, sentence length, level of formality and preferred terminology.

That is useful training material in the ordinary sense of the word: evidence of how the business communicates.

I would still be careful about copying marketing language too aggressively.

Nobody wants a support response that sounds like the hero section of a landing page.

Support should sound natural.

The company’s language gives it a home to sound natural inside.

Tiny words can carry operational meaning.

Pricing pages are particularly fun for this.

Words such as: per user, per workspace, included, starting at, up to, additional and fair use can completely change an answer.

Take: “100 conversations included each month.”

I immediately have questions.

Does the allowance reset on the calendar month or billing date?

What exactly counts as one conversation?

What happens at 101?

Can unused conversations roll over?

Is the allowance shared across a team?

Does a reopened chat count again?

One pricing sentence can generate half a support article.

The site teaches the system where to start.

It can also reveal where the public explanation stops.

Then the customer asks something the website cannot possibly know.

This is where the limits become obvious.

A customer asks: “Has my order shipped?”

The website knows the shipping policy.

It does not know whether this order shipped.

Another asks: “Why was my payment declined?”

The public site might explain accepted payment methods.

It has no idea what happened to that transaction.

Someone asks: “How many conversations do I have left this month?”

The pricing page explains allowances.

The answer requires account state.

This is the line I care about.

Public knowledge can get us surprisingly far.

Personal questions require personal context.

The website can explain the rules of the game. It usually cannot tell you what is happening to this particular player right now.

Account state changes ordinary answers.

Suppose a customer asks: “Can I invite another teammate?”

The website says the Business plan supports five team members.

Great.

Now suppose the customer is already on Business and already has five members.

The generic answer: “Yes, Business supports up to five team members” is technically related and practically useless.

Their real situation is that the workspace has reached its limit.

A useful response might explain the current member count, what needs to happen before another invitation can be sent and whether an upgrade is available.

The public rule stayed the same.

Context changed the answer.

Orders are another perfect example.

E-commerce makes this distinction incredibly easy to see.

A shop can publish: “Orders usually ship within two business days.”

Customer: “When will mine ship?”

Repeating “usually within two business days” may be the best possible answer if no order access exists.

Once the support system can see that the order was placed yesterday, payment cleared and fulfilment is currently marked as packing, the response gets much better.

Add carrier information later and it gets better again.

This is how support grows beyond website knowledge.

You start with general truth.

Then add situational truth.

The two belong together.

Some important knowledge should never be public.

There is another category the website should not know because the website is public.

Internal procedures.

Fraud rules.

Escalation contacts.

Staff instructions.

Exceptions that employees may offer under specific circumstances.

A business might publicly say: “Refunds are available within fourteen days.”

Internally, support may have permission to make an exception when a service outage affected the customer.

That does not belong on the public pricing page.

It may still belong in the support system.

This is where private knowledge starts becoming genuinely valuable.

It fills gaps that the website was never supposed to fill.

The difficult part is deciding what the system may actually say.

Knowing something and being allowed to reveal it are separate concerns.

Imagine the system has access to an internal note: “Customer has been flagged for manual review.”

The customer asks: “Why is my withdrawal delayed?”

We definitely do not want the response to casually expose every internal field available to it.

The same problem appears with discounts.

Internal note: “Retention team can offer up to 20%.”

That information may help a human agent.

It does not mean the automated support agent should announce it to every person who asks about cancellation.

Knowledge needs boundaries.

Public.

Private but answerable.

Private and operational.

Human-only.

Those distinctions matter.

I would rather have a smaller trustworthy brain than an enormous messy one.

There is a very tempting idea in AI products: upload everything.

Every PDF.

Every old document.

Every Slack export.

Every policy.

Every support transcript.

More knowledge sounds automatically better.

I do not think it works that cleanly.

Imagine giving a new employee a folder containing every document the company has produced since 2019 without explaining which ones still apply.

They would know a lot.

They would also be dangerous.

Old prices.

Cancelled features.

Draft policies.

Internal speculation.

Duplicate instructions.

Contextless screenshots.

The problem is not a lack of text.

The problem is deciding what deserves trust.

A knowledge system becomes more useful when the business can explain which information is authoritative, which is contextual and which should disappear. Uploading another thousand pages is much less exciting once contradictory page number 847 answers a customer.

Clean knowledge also makes mistakes easier to investigate.

Suppose the support system gives a bad answer.

“Your plan includes unlimited team members.”

It does not.

Someone now needs to understand why that happened.

Was the pricing page wrong?

Was an old article still indexed?

Did an internal note override something newer?

Did the system misunderstand the question?

If the knowledge has structure, this can be investigated.

If the system received a giant unsorted dump of company text, debugging becomes archaeological work.

That is one reason source visibility matters so much.

I want to be able to look at a strange answer and trace where it came from.

Otherwise improving the system turns into guessing.

Websites also expose questions that should have been answered earlier.

There is a funny side effect to connecting support closely with website knowledge.

You start noticing where the website itself could improve.

Suppose twenty people ask: “Does this work on WordPress?”

The answer is yes.

The support system handles it perfectly every time.

Great.

Eventually somebody should probably put “Works with WordPress” somewhere obvious.

Support should absorb uncertainty.

It should also reveal where uncertainty keeps appearing.

Repeated conversations become feedback about the public site.

I really like that loop.

Website teaches support.

Support teaches the website.

The information gets better from both sides.

A strange question can reveal missing business knowledge.

Sometimes a customer asks something nobody has documented anywhere.

“Can I transfer my account to another company after an acquisition?”

Search the website.

Nothing.

Search internal notes.

Nothing.

Ask the team.

Three people have three different opinions.

At this point support has discovered something useful: the business itself has not decided the answer clearly enough.

I think this is one of the underrated benefits of support work.

Customers are excellent at finding corners of a business nobody remembered to formalise.

Once a decision is made, that answer can become knowledge.

The next customer does not need the same internal debate.

Good knowledge grows from real questions.

I am suspicious of giant knowledge-base projects created entirely in advance.

You can spend weeks imagining everything customers might ask.

Then customers arrive and ask something completely different.

Real conversations are better teachers.

Start with the website.

Add the internal material you already know matters.

Then watch what people actually ask.

A question appears repeatedly.

Add a clean answer.

An existing page causes confusion.

Fix the page.

An edge case appears.

Decide how it should work.

Knowledge improves through use.

That feels much healthier than trying to predict the entire future of customer curiosity before anyone has asked anything.

There is a difference between reading a page and understanding its role.

Consider a page titled: “Enterprise Security”.

It might say: “Single sign-on is available for enterprise customers.”

Straightforward.

Now a customer on the cheapest plan asks: “How do I enable SSO?”

The system needs several pieces of understanding.

SSO exists.

It is restricted to Enterprise.

This customer is not on Enterprise.

Their current question cannot be solved by giving them setup instructions.

The useful answer explains availability first.

Maybe it tells them where to upgrade or who to contact.

This sounds obvious when a human reads it.

Software needs all the pieces in the right relationship.

Product context can be as simple as knowing the page.

Context does not always require private account data.

Sometimes knowing what page the visitor is looking at is already useful.

A person asks: “Does this include setup?”

If they are on the Business pricing card, the meaning may be obvious.

If they are on an implementation guide, they could be asking something completely different.

Page context can help interpret short questions.

So can the product they are viewing.

So can the plan currently selected.

So can the previous two messages.

Small context adds up quickly.

The website is especially good at giving support a starting vocabulary.

One of the first things I would want any support system to understand is the nouns of the business.

What is a workspace?

What is a conversation?

What is a project?

What is a booking?

What is a seat?

What is a top-up?

These words can sound generic while having very specific meanings inside a product.

If a customer says “account” while the product internally distinguishes between user, workspace and organisation, support needs to understand the ambiguity.

A website usually contains enough examples to begin mapping that vocabulary.

This matters because language is where many misunderstandings begin.

There are also things a website says accidentally.

This part is fun.

Sometimes the structure of a website reveals operational reality the copy never states directly.

There are twelve different setup guides.

That probably means implementation varies by platform.

There is an unusually detailed migration page.

Migration is probably a common concern.

The refund FAQ is six paragraphs long.

Refunds probably contain exceptions.

Every pricing card mentions team members.

Team access is clearly central to the product.

Support can benefit from these clues even when they are not formal policies.

They help us understand where the business expects questions to cluster.

Then there is the problem of things the business forgot existed.

Old landing pages are incredible.

You think a plan disappeared two years ago.

Google finds: /summer-2024-pro-plan.

It is still live.

It still says $29.

Current pricing says $49.

The support system crawls both.

Now we have invented a new customer service problem.

This is why automatic website research needs controls.

Discovery is great.

Businesses still need a way to exclude pages, correct sources and tell the system what should win.

Automation can find the mess incredibly quickly.

It should not pretend the mess is truth.

I think the first setup experience should respect what already exists.

This idea has influenced how I think about onboarding.

Imagine buying a support product and immediately seeing an empty dashboard.

“Add your knowledge.”

Okay.

Where do I begin?

Copy the homepage?

Upload pricing?

Write FAQs?

Explain every feature?

Suddenly I have another project.

I much prefer the opposite feeling.

Give the system the website.

Let it do some homework.

Come back with a useful starting point.

Then show me where I should add the things it could never learn publicly.

That feels respectful of the work the business already did.

It also makes the first useful result arrive much earlier.

Thirty seconds of research can save hours of boring setup.

There are plenty of tasks where manual configuration is valuable.

Copying publicly available facts into another dashboard is not a particularly exciting use of anyone’s afternoon.

Business name.

Product names.

Pricing.

Public policies.

Basic implementation instructions.

Existing FAQs.

If the system can collect these itself, let it.

Save human attention for the things that require human judgment.

Which sources matter most?

What should remain private?

What exceptions exist?

When should a human take over?

Those are much better questions to spend setup time answering.

The website should remain a living source.

Another mistake would be treating website research as a one-time import.

Websites change.

Pricing changes.

Features change.

New articles appear.

Policies get rewritten.

A support system trained on a snapshot eventually becomes a historian.

Sometimes that is useful.

Usually customers want the current product.

Keeping website-derived knowledge fresh should therefore be part of the relationship.

How often depends on the business.

A restaurant menu might change weekly.

A legal policy might change twice a year.

A software documentation site can change every day.

Freshness is part of correctness.

Corrections need to be easier than fighting the website.

Imagine the website says: “Setup takes five minutes.”

The business knows enterprise customers usually require an additional approval step.

They should be able to teach that nuance without rewriting the whole website for one edge case.

Public content can remain broad.

Support knowledge can be more specific.

That is an important relationship.

The system should be able to learn: “For normal installations, use the public five-minute guidance. For enterprise accounts with restricted deployment, mention the approval step.”

Now the business has added useful context without duplicating everything else.

This is where I stop wanting the system to be clever.

There is a point where “smart” behaviour becomes uncomfortable.

If two sources disagree and there is no clear authority, I do not want the system creatively resolving the disagreement.

If a customer-specific question requires data the system cannot access, I do not want an educated guess.

If a policy has an exception nobody documented, I do not want improvisation.

There are moments where the most intelligent behaviour is recognising the edge of available knowledge.

“I can confirm the standard policy, but I cannot see the status of your order.”

Fine.

“I found two different published limits and I do not want to give you the wrong one.”

Also fine.

Confidence should come from evidence.

A website can give you a surprisingly capable first version.

After spending so much time on the limits, I do not want to undersell the beginning.

A good website can teach a support system an enormous amount.

Enough to answer what the product does.

Enough to explain plans.

Enough to describe installation.

Enough to explain public policies.

Enough to understand common terminology.

Enough to recognise many of the questions customers are likely to ask before they ever create an account.

That is a strong first day.

Especially when the alternative is an empty support system waiting for someone to manually rewrite all of that information.

Then the second layer makes it belong to the business.

The interesting work starts after the automatic part.

Add the internal rules.

Add the operational details.

Connect the customer context that is genuinely useful.

Decide which information may be revealed.

Decide where automation should stop.

Review the strange conversations.

Correct knowledge as the product changes.

Add answers to questions nobody predicted during setup.

Over time the system starts knowing the business in a way a crawl alone never could.

That is healthy.

The crawl gives it a head start.

Experience gives it depth.

I think this is the useful middle ground.

There are two extremes that both feel strange to me.

One asks a business to manually explain everything before the support system can answer anything.

The other points at a website and acts as if every future customer question has now been solved forever.

Real businesses are messier and more interesting than either version.

Public pages contain a huge amount of useful knowledge.

Private operations contain more.

Customer context changes some answers.

Real conversations expose missing pieces.

Humans remain valuable when uncertainty gets too high.

Put those pieces together and support begins to feel less like a chatbot attached to a website and more like a system that actually understands where the answer should come from.

The website is an excellent first teacher. The mistake would be assuming it is the last one.

The part I find most exciting is how little work the beginning can require.

A business has already spent years teaching the internet who it is.

Its products are public.

Its language is public.

Its pricing may be public.

Its FAQs are public.

Its documentation may contain hundreds of detailed answers.

There is no reason to pretend all of that work starts from zero when support is installed.

Read it.

Organise it.

Understand what appears current.

Notice where the gaps are.

Then ask the business for the information only the business can provide.

I think that is a much nicer relationship between automation and setup.

The machine does the boring discovery.

The human supplies judgment.

Customers get useful answers earlier.

And when the website finally runs out of things to teach, that is not a failure.

It is simply the point where the real business begins.