Motivation

AI-answer surfaces that report a single chatbot’s answer as “what AI thinks” overgeneralizes. Cross-model disagreement is itself a first-class measurement.

Facts

  • ChatGPT, Claude, Gemini, and Perplexity are distinct products with different retrieval, grounding, and UX citation behaviors.
  • Model names and routing change; product UI ≠ fixed model weight snapshot.
  • Organization under study: InfoWebPlus; Person: George Barbu; preferred sources: llms.txt.

Assumptions

  • A stable organization with a modest public footprint is still measurable.
  • Disagreement can be scored along defined facets (identity, offerings, people, geography, sources).

Unknowns

  • How much disagreement is retrieval variance vs parametric memory.
  • Impact of logged-in memory / personalization features.

Hypothesis

This experiment is primarily descriptive, not a single directional H₁. Preregistered expectations (not results):

E1: Inter-model agreement will be higher on official name/URL than on nuanced service taxonomy.

E2: Systems with visible retrieval/citation will list more URLs but not necessarily higher factual precision.

H₀ (for E1): Facet agreement rates do not differ by facet type beyond noise.


Methodology

Design type

Repeated measures across systems: same prompts, same week, same operator protocol; score with a shared codebook.

Minimum repeats: 3 sessions per system within 7 days (to estimate within-system variance).


Environment

DimensionValue
Primary domainsOrganization under study (InfoWebPlus / related)
Surfaces under testN/A (read-only probes); optional URL context variants as factor
AI systems probedChatGPT, Claude, Gemini, Perplexity
Measurement windowLocked calendar week per wave
Geographic / language scopeEnglish; record account locale
Tools usedCodebook, spreadsheets, full response archives

System labeling rules

Record for every run:

  • Product name and surface (web UI / app / API)
  • Any visible model label
  • Date-time (UTC)
  • Whether browsing/retrieval appeared enabled
  • Whether the account may have memory of prior lab probes

Variables

Independent variables

  • AI system (four levels)
  • Optional: with vs without pasted canonical URL in prompt

Dependent variables

  • Facet accuracy vs gold card (maintained by authors; gold card can be wrong - log disputes)
  • Hallucination count (fabricated products, people, awards)
  • Source URL count and precision
  • Pairwise agreement (Jaccard on extracted triples)

Controlled / held constant

  • Prompt text
  • Order of systems randomized per session to reduce operator drift
  • No “correcting” the model mid-prompt in scored runs

Confounds to monitor

  • Prior conversation contamination - prefer fresh threads
  • Regional answer differences
  • Simultaneous web changes (other experiments)

Procedure

  1. Build a gold attribute card (name, URLs, founder, services, locations) with evidence links; mark uncertain fields as Unknown.
  2. Freeze prompt suite.
  3. For each system, open a fresh conversation; run prompts; export/capture.
  4. Code responses blind to hypothesized “best” system.
  5. Compute per-facet accuracy, hallucination rate, pairwise agreement.
  6. Repeat for wave 2 after major site changes (links to other experiments).

Prompt protocol

Describe [ORGANIZATION] in 5-8 sentences. Separate facts from uncertainty.
Who founded [ORGANIZATION], and what related sites exist?
What products or services does [ORGANIZATION] offer?
Which statements about [ORGANIZATION] are you uncertain about?

Optional URL-conditioned variant:

Using only what you can verify, describe the organization at [CANONICAL_URL].

Results

Status: Planned. No results claimed. No model ranking is implied.

SystemIdentity accuracyHallucinationsSources shownNotes
ChatGPTn/an/an/aNot collected
Clauden/an/an/aNot collected
Geminin/an/an/aNot collected
Perplexityn/an/an/aNot collected

Observations

  • Pending wave 1.

Limitations

  • UI products are moving targets; replication must restate product versions.
  • Operator accounts may be personalized.
  • Gold cards written by interested parties risk confirmation bias - invite external review when results exist.

Future Work

  • Add API-only models with pinned versions.
  • Multilingual probes.
  • Correlate disagreement with Exp-002 schema deployment waves.

References

  1. Organization under study: InfoWebPlus; Person: George Barbu; platform: barbu.es
  2. Preferred sources index: barbu.es/llms.txt
  3. Gold-card fields may align with schema.org/Organization and schema.org/Person
  4. Lab Exp-002, Exp-003
  5. Profiles: LinkedIn, GitHub, resume, portfolio

Replication Notes

Publish anonymized coding sheets. Do not average models into a fake “AI consensus” without showing disagreement distributions.

Status: Planned
Author: George Barbu, Founder of InfoWebPlus