Skip to content
CiteSprout

CiteSprout

Replayable, transcript-backed GEO scoring

Created on 8th August 2026

•

CiteSprout

CiteSprout

Replayable, transcript-backed GEO scoring

What is the problem your project solves?

CiteSprout solves for Generative Engine Optimization in verifiable, retracable and transparent way.

If Claude or ChatGPT never names your brand in the answer, you are invisible to that buyer. There is no analytics tab for this. The tools that do measure it hand you a score with nothing behind it, so you cannot tell whether a low number means "the model dislikes us," "the model has never heard of us," or "the tool made it up."

CiteSprout answers the question with evidence: 15-20 realistic buyer prompts, fired at Claude and GPT with web search enabled, every raw response persisted before anything is derived from it. Click any number in the viewer and you get the response text, the model snapshot that produced it, the citations it returned, and whether the call needed a retry.

How you are solving it?

CiteSprout answers the visibility question with evidence instead of a bare number.

  • Prompt matrix: generates realistic buyer prompts (informational, transactional, navigational) for a brand's niche and competitors, weighted to match GEO-bench's real-world query distribution (Aggarwal et al., KDD 2024).
  • Orchestrator: fires every prompt at Claude and GPT with web search/grounding enabled, with retry-then-FAILED handling so one bad call never blocks the run.
  • Parser: extracts brand mentions, sentiment (positive/neutral/negative), citations, and competitor mentions per response.
  • Classifiers: prompt-category and sentiment classification were designed around a ColBERT-style late-interaction matcher and a hosted RoBERTa sentiment model. What ships is a single-vector embedding classifier (Hugging Face or OpenAI) standing in for the ColBERT-style matcher, with RoBERTa used for sentiment when HF_API_KEY is set and a Claude-based classifier as the fallback otherwise — the code reports which backend actually ran rather than assuming one.
  • Store: persists every raw response verbatim to SQLite the moment it returns — nothing derived is kept without the transcript behind it.
  • Score engine: a counted-mention rate per intent category, combined as a weighted average (80/10/10), with categories under 3 successful records reported as "insufficient data" rather than a misleading rate.
  • Viewer: the score, a per-category breakdown, a diagnostics tab for failed/unparseable calls, and a methodology explainer — click any number to the exact model response it came from.

All of this was built during the hackathon window. The plan/architecture docs were drafted first, then implemented end-to-end against real Claude and GPT API calls. No prior-hackathon code was reused.

How Did You Use Claude?

Claude is a core part of the intelligence layer, not a bolt-on:

  • Grounded provider calls: Claude (with web search) is one of the two LLMs actually being measured for brand mentions — it's both a subject of the audit and, via the Anthropic SDK, the audited caller.
  • Prompt generation: the prompt matrix engine uses Claude to expand a small set of hand-written seed prompts into the full realistic buyer-prompt matrix.
  • Fallback sentiment classifier: when the optional Hugging Face RoBERTa backend isn't configured, sentiment classification of "was the brand mentioned favorably?" falls back to a Claude-based classifier.
  • Suggestions engine: paper-backed improvement suggestions are derived from each run's own transcripts using Claude to reason over the parsed evidence, not a canned template.

What is the deployed URL for this project?

https://citesprout-geo.fly.dev/

Cheer Project

Cheering for a project means supporting a project you like with as little as 0.0025 ETH. Right now, you can Cheer using ETH on Arbitrum, Optimism and Base.

Discussion

Builders also viewed

See more projects on Devfolio