Human-AI
Overview
Section titled “Overview”Strips mechanical writing signals from English text and restores rhythm, precision, and voice. Combines pattern detection (43 patterns across 3 tiers), statistical rhythm measurement (burstiness, TTR, entropy), and voice injection into a single iterative skill with scoring.
The goal is a better text, not a fooled detector: no rewrite can guarantee that a tool will classify the result as human, and AI-detector scores are not a quality criterion — those tools misfire often and penalize neurodivergent and non-native writers disproportionately.
The statistical metrics serve as measurable proxies for natural rhythm, not as a scoreboard to beat. The research behind the skill is useful for a different reason: it shows that swapping synonyms changes nothing, because the problem lives in the sentence’s architecture — length, rhythm, clause structure, and information density.
Companion to the humanizar skill (PT-BR). For Portuguese text, use humanizar.
When to Use
Section titled “When to Use”- English text sounds generic, bland, or AI-generated
- Requests like “humanize”, “de-slop”, “remove AI patterns”, “make it sound human”
- Rewriting with voice and personality, or fixing a generic tone
- Agent pipelines generating English content
- Review another agent’s output before publishing
What Makes It Different
Section titled “What Makes It Different”| Feature | blader/humanizer | brandonwise/humanizer | human-ai |
|---|---|---|---|
| Patterns | 29 | 28 + statistical | 43 (incl. P31-P43 emerging 2026) |
| Iterative scoring | ❌ | ❌ | ✅ 0-100 with loop |
| Metrics script | ❌ | ❌ | ✅ scripts/measure.py |
| Voice presets | ❌ | ❌ | 7 presets |
| Anti-synonym-swap | ❌ | ❌ | ✅ Structural rewrite enforced |
| Empirical baselines | ❌ | Partial | ✅ SSRN + GPTZero + NeurIPS |
Operation Modes
Section titled “Operation Modes”| Mode | When to use | Result |
|---|---|---|
full_mode |
Default — “humanize this” | Metrics + diagnosis + rewrite + scoring |
direct_mode |
Pipelines or “humanize quick” | Final version + synthetic report |
review_mode |
Audit text from another agent | Aggressive rewrite + before/after metrics |
Installation
Section titled “Installation”npx skills add https://github.com/fabricioctelles/skills --skill human-aiResearch Foundation
Section titled “Research Foundation”| Source | Finding |
|---|---|
| RAID Benchmark (ACL 2024) | Structural paraphrasing drops DetectGPT from 70.3% → 4.6% |
| humanizerai.com (2026) | Vocabulary bans HURT performance by 43 percentage points |
| GPTZero | Burstiness (sentence length variance) is the primary detection signal |
| SSRN | Human TTR: 0.553 vs AI: 0.455 |
| NeurIPS 2023 | Human intrinsic dimensionality ~9 vs AI ~7.5 |
| Washington Post | “It’s not X, it’s Y” = #1 AI tell across 328K messages |
Part of this research was produced by measuring detector bypass rates. The skill uses those findings for what they reveal about rhythm and structure — not as a target. None of these numbers proves human authorship, and none of them decides what gets rewritten.
Changelog
Section titled “Changelog”v1.0.1 (Aug 2026)
- Positioning aligned with
humanizar: the skill no longer promises “undetectable” text, and AI-detector scores are no longer a criterion - Statistical metrics repositioned as proxies for natural rhythm, not a scoreboard
- Dropped the contraindication based on “text already validated as human by multiple detectors”
v1.0 (Jul 2026)
- Initial release with 43 patterns (P1-P43), including 13 emerging from 2026
- 3 operation modes with iterative scoring (0-100) and strategy fallback
- 7 voice presets (Essay, Journalistic, Academic, Corporate, Social, Casual, Legal, Instructional)
scripts/measure.pyfor deterministic metric calculation (zero dependencies)- 7 documented gotchas from real operational failures
- Empirical baselines calibrated from published research
- Bidirectional cross-reference with
humanizarskill (PT-BR) - Based on blader/humanizer, brandonwise/humanizer, and Aboudjem/humanizer-skill
📄 Full documentation on GitHub
TypeSafe Jev Integration
Section titled “TypeSafe Jev Integration”This skill supports optional integration with TypeSafe Jev — a judgment engine providing calibrated probability assessments for subjective criteria.
When to Use Jev
Section titled “When to Use Jev”| Evaluation | Use Jev? | Method |
|---|---|---|
| Pattern detection | No | Deterministic (regex, patterns) |
| Pattern severity | Yes | Score (0.0-1.0) |
| Voice match quality | Yes | Score |
| Rewrite confidence | Yes | Score |
| Change classification | Yes | Noul (structural/cosmetic/hybrid) |
Available Questions
Section titled “Available Questions”- 8 Score questions: pattern severity, voice adequacy, fluency, tone preservation, naturalness, restored entropy, achieved burstiness, overall confidence
- 12 Noul questions: change type, pattern category, intervention level, factual preservation
Discovery Protocol
Section titled “Discovery Protocol”from typesafe import jev_available
if jev_available(): from typesafe import Score, Noul # Use Jev for subjective criteriaelse: # Fall back to heuristic scoringWhen Jev is unavailable, the skill uses internal heuristic scoring — functional but less nuanced.