https://sharaj.pages.dev/ · Developer Tools
4 items
Test, analyze, and version system prompts against the LLM you ship on. BYOK, local-first, no backend. litmus turns "does this prompt work?" from a gut feeling into a measured result — for plain prompts, tool calls, and multi-step agents alike. Paste a system prompt, pick the model you actually ship on, and choose what you're testing. Everything runs locally in your browser with your own API keys. There is no litmus backend, no account, and no tracking. ── TWO WAYS TO TEST ── 1) OUTPUT QUALITY — litmus analyzes your prompt, auto-writes a rigorous LLM-as-judge rubric per quality dimension, generates typical/edge/adversarial test cases, runs them on your target model, and scores each output. Then it proposes ranked fixes and can auto-apply them for the next pass. 2) TOOL & AGENT BEHAVIOR — define your tools (JSON schema) and litmus checks, deterministically (no LLM judge), that the model calls the right tool with valid arguments and avoids the ones it shouldn't. For agents, define a goal plus mock tools with scripted results (inject a failure to test recovery); litmus runs the model in a multi-step loop and scores the trajectory across goal completion, tool selection, argument validity, recovery, and efficiency. This mode skips the rubric steps — pick it on the first screen and go straight to your tests. ── WHAT YOU GET ── • Auto-generated rubrics and test cases — including tool tests proposed from your catalog. • Deterministic tool/agent checks that don't drift run-to-run. • Variance built in — run each case N times to see the spread (mean ± range), so a noisy score is visible, not hidden. • Speed measured live (time-to-first-byte, tokens/sec) for quality runs. • Versioning — every run is saved; reload any version, compare by dimension, export as Markdown or JSON. • Works with OpenAI, Anthropic, and Google targets. ── PRIVACY & CONTROL ── • Bring your own key (BYOK). Keys are stored only in your browser. • Local-first — no litmus servers. Your data goes only to the provider you choose, to run the test. Tools in agent runs are mocked — nothing real is executed. • No analytics, no ads, no account. • A spend cap you set blocks runs that would cost more than you want. ── GOOD FOR ── Prompt engineers and AI app developers who want to quickly verify a prompt, tool, or agent before shipping — without standing up a cloud eval platform. Pick a judge model different from your target to reduce self-preference bias and get more trustworthy quality scores.
Aug 2, 2026
rating_count is the Chrome Web Store ratings count, not a written-review count.
Media assets
Screenshots and videos on the listing.
Has promo video
Whether the listing includes at least one video.
Languages
Declared language locales.
Developer website
Listing exposes a developer website URL.
Contact email
Listing exposes a contact email.
Keyword in name
Case-insensitive substring match in the name.
Keyword in description
Case-insensitive substring match in the description.
Keyword occurrences in description
Count of case-insensitive occurrences in the description.
Category user-count percentile
Share of same-category extensions with fewer users (null if unknown).
These are transparent listing completeness / keyword signals, not a prediction of Chrome Web Store search ranking.