Arooj Review
How Reviews Are Built

The testing methodology behind
every review on this site

This page documents exactly how tools are researched, tested, scored, and written about. If you want to understand why a verdict is what it is - or verify that a finding is grounded in evidence - this is where to start.

The Review Process

Seven steps from tool access to published verdict

No step is skipped. No review is published before the process is complete.

01
Multi-platform research before any testing begins
Before signing up for a tool or opening a browser tab, structured research is conducted across search engines, AI synthesis tools, community platforms, and SEO intelligence sources. This establishes what is known, what is disputed, and what needs to be verified directly.
Platforms consulted: Google, Bing, DuckDuckGo · ChatGPT, Claude, Gemini, Grok · Reddit, Quora · Ahrefs, Semrush · Trustpilot, G2, Capterra, Product Hunt
02
Semantic content mapping - the article plan before any writing
Inspired by Koray Tugberk Gubur's Holistic SEO framework, a semantic content map is built before writing begins. This maps primary keywords, search intent, named entities, attribute clusters, PAA questions, competitive gaps, and freshness signals. The article is never written before this map exists.
Map elements: Primary keyword · Search intent classification · Target audience segment · Secondary keywords · Named entities · Attribute clusters · PAA + related queries · Sub-topics · Competitive gap analysis · Freshness signals
03
Direct hands-on testing - every publicly accessible feature
The tool is signed up for and tested under the conditions a typical user would encounter - free tier first, then paid features where accessible. Every feature described in the review is tested by completing an actual task, not by reading a tooltip or a marketing page. What works, what doesn't, and what behaves differently than advertised is documented during testing.
Testing protocol: Sign up under typical user conditions · Test each feature by completing a real task · Document friction points, performance observations, and UI/UX findings · Record what works, what doesn't, and what contradicts vendor claims
04
Cross-verification of findings against independent sources
After forming a personal analysis from testing, findings are cross-referenced against user reviews, independent studies, YouTube walkthroughs from non-sponsored creators, and competitor comparisons from credible publications. Where independent research contradicts or qualifies the testing findings, both are presented.
Sources consulted: G2, Capterra, Trustpilot, Product Hunt · Reddit threads from verified users · Independent accuracy studies · YouTube walkthroughs from non-sponsored creators · Credible SEO and tech publications
05
Pricing verification - on the official plan page, on the publication date
Every price cited in a review is verified against the official plan page on the date of publication. The publication date is recorded alongside the pricing data. Grammarly, for example, changed its plan structure four times between 2024 and 2026 - a review that doesn't note the verification date gives readers false confidence about current pricing.
Verification standard: Official plan page only - never from third-party aggregators · Verification date recorded · Readers explicitly directed to verify before purchasing · Plan name changes, feature consolidations, and discontinuations noted
06
Honest limitation disclosure - with equal weight to strengths
Limitations are not footnotes. Each limitation is stated factually, contextualized (is this a fundamental product gap, a free-tier restriction, a known bug, or an early-stage feature?), assessed for impact (who does this affect and how significantly?), and given an improvement direction. Omitting a limitation because a brand is popular or a commission is involved is not something that happens here.
Limitation format: State the limitation factually · Provide context · Acknowledge the impact · Suggest what the improvement would look like · Never call a tool broken without context · Never give a pass on a serious flaw because the brand is popular
07
Writing - built for the human reader, optimised for search as a byproduct
Every piece is written following the precision standard of the Ahrefs Blog (no filler, action before explanation, data cited not implied), the E-E-A-T signal principles of Lily Ray (experience-first credibility, named sources, transparent limitations), and the topical completeness requirement of Koray's framework (no reader should need a second tab to answer a follow-up question the content naturally raises). SEO optimisation is a byproduct of genuine quality - not the other way around.
Simultaneous optimisation targets: SEO (traditional search) · GEO (generative engine optimisation for ChatGPT, Gemini, Perplexity) · AEO (answer engine optimisation for featured snippets) · LLMO (large language model optimisation for accurate citation by AI tools)
Research Sources

Where findings are verified

No single source is treated as definitive. Cross-verification across multiple independent platforms is required before a claim is stated as fact.

🔍
Google, Bing, DuckDuckGo
Search and discovery - finding what is publicly known and disputed about a tool
G2, Capterra, Trustpilot
Verified user reviews - real user experiences including edge cases and long-term usage
💬
Reddit, Quora
Community intelligence - real user feedback, pain points, and workarounds
🧠
ChatGPT, Claude, Gemini, Grok
AI synthesis and perspective - cross-checking findings against AI knowledge bases
📊
Ahrefs, Semrush
SEO intelligence - keyword data, topical gaps, SERP analysis, competitive landscape
📑
Official documentation
Feature specifications, support pages, privacy policies, security certifications
🎬
YouTube walkthroughs
Non-sponsored independent creators showing real workflows and edge cases
🔬
Independent research studies
Peer-reviewed or methodologically documented studies on tool accuracy and performance
Scoring System

What the scores mean and how they are assigned

Every score is Arooj's assessment based on direct testing. Scores reflect the tool's performance for its intended primary use case - not overall impressiveness.

Scores are assigned on a 1 - 10 scale and reflect how well the tool does the specific job it is positioned to do. A grammar checker is scored on grammar checking accuracy, not on whether it has a good generative AI feature. A plagiarism checker is scored on detection reliability and source coverage, not on how nice the interface looks.

The score reflects the primary use case. Secondary features that work well can lift the score. Secondary features that are overstated in marketing - and don't hold up to testing - can lower it. The Grammarly score of 7.5 reflects strong core proofreading and genuine integration value, offset by AI detection claims that independent testing does not support at the level marketed.

Scores are updated when tools improve significantly, when pricing changes affect value-for-money, or when independent research changes the evidence base for a specific claim.

Score range What it means Label
9.0 - 10.0 Best in category. Does what it promises at the claimed level. Limitations are minor and clearly disclosed. Excellent
7.5 - 8.9 Strong for its primary use case. Some secondary features may be overstated. Worth the cost for the right user. Recommended
6.0 - 7.4 Useful within a narrower scope than marketed. Meaningful limitations exist. Best for a specific user type. Conditional
4.0 - 5.9 Works in limited scenarios. Significant gaps between marketing claims and real-world performance. Limited
Below 4.0 Does not reliably do what it claims. Not recommended at current pricing or feature level. Not Recommended
Testing Matrix

What is tested in each review category

The testing criteria differ by category because the use cases differ. What matters for a grammar checker is not what matters for an AI video generator.

Category Core feature test Pricing verified User reviews consulted Independent studies Free tier tested
Writing & Editing Error detection accuracy vs Word and Google Docs Where available
SEO & Content Content scoring consistency, keyword data accuracy
AI Detection False positive rate on human text, detection on edited AI ✓ Required
AI Assistants & LLMs Output quality, reasoning, context retention, voice fidelity
AI Image Generation Prompt fidelity, output consistency, style control Where available
AI Video Generation Output quality, coherence, lip sync (where applicable) Where available
AI Audio & Podcast Voice naturalness, editing accuracy, export quality Where available
Plagiarism Checkers Detection rate on known content, false positive rate ✓ Required
Limitations Policy

How testing gaps and access limitations are handled

Not every feature can be tested on every plan. This policy governs how untestable features are handled - transparently, not silently.

01

Paywalled features that were not tested

Stated explicitly: "I was not able to test [Feature X] during this review because [specific reason]. Based on available documentation and third-party user reports, here is what is reported about it - but treat this as unverified from my direct experience." No fabrication. No filler description from a marketing page.

02

Conflicting independent research

When independent studies conflict - as they do on AI detector accuracy, where results range from 22% to 99% depending on the study - both findings are presented with methodology context. The most favorable study is not chosen to match a preferred conclusion. The conflict is the finding.

03

Time-sensitive data and model updates

Freshness signals are tracked. Claims about tool performance that are time-sensitive - AI model accuracy, pricing, feature availability - are flagged for future verification and dated in the review. Static facts don't expire; dynamic data is tracked.

04

Regional and access limitations

Features that behave differently by region, account type, or plan tier are noted explicitly. A feature available in the US Enterprise plan that isn't available in the UK free tier is not described as universally available. Access conditions are part of the review.

Questions or Corrections

Found something that needs updating?
Want to suggest a tool for review?

If you have found a factual error, a pricing discrepancy, or a finding that newer evidence contradicts - the contact page is where to send it. Every correction is taken seriously and reviewed against the source.