Why this is being tested

Document the method and publish dated observations as they accumulate.

Published with the method before the result, because a conclusion announced without a stated method, on a sample nobody can see, is an opinion with a number attached.

Result

No result yet. Baseline task battery defined and first runs scheduled; observations will be published here with the scoring sheet.

Hypothesis

Agents will extract entities and services accurately from semantically structured sites and fail in predictable ways on sites where content lives in scripts, images or hover states.

Method

  • Select a small sample of real business websites across categories, including this one.
  • Run the same task battery against each with a browser agent: identify the legal entity, list services, find pricing or engagement terms, locate a contact action and describe how to complete it.
  • Score outputs against ground truth established by manual review.
  • Record the failure type when extraction fails: missing content, ambiguous labels, inaccessible controls, or hallucinated filler.

Environment. Browser agent over live public sites; manual ground-truth review

Limitations

  • Small sample; results describe failure patterns, not category-level rates.
  • Agent behavior changes with model versions; runs are dated and versioned.

Replicating this

Replicable with any browser agent: fix the task battery before looking at any site, score against manually established ground truth, and date the model version used.

What this means for an operator

Open your own site with images and scripts disabled and see which of your key facts survive.