I will professionally evaluate AI-generated content and LLM responses in Persian (Farsi) and English.
As a PhD researcher in Cognitive Sciences with extensive experience in research, content analysis, evaluation frameworks, and structured assessment, I can review AI outputs beyond basic grammar checking.
I can evaluate responses for:
• Accuracy and factual consistency
• Relevance to the user’s instruction
• Instruction following
• Logical coherence
• Clarity and readability
• Persian language fluency and naturalness
• Tone and cultural appropriateness
• Hallucinations and unsupported claims
• Response quality and completeness
• Safety and content quality
• Ranking and pairwise comparison of AI responses
I can also work with your existing evaluation rubric or help develop a clear scoring framework for your project.
Typical projects include:
• LLM response evaluation
• AI training datasets
• Human feedback / preference evaluation
• Persian localization QA
• Prompt-response review
• Annotation quality review
• AI-generated content auditing
• Golden-response creation
• Fact-checking and research-based evaluation
Deliverables can be provided in Excel, CSV, Google Sheets, Word, or your preferred annotation format.
For larger datasets or ongoing AI evaluation projects, please contact me for a custom offer.