Universal Design AI Systems: Automating WCAG Audits & Fixing Violations with AI Agents

The Vision

Universal Design (universell utforming) means building digital products that work for everyone. In Norway, it's the law. But manual accessibility audits don't scale β€” 400+ municipalities, each with dozens of websites.

Two projects solve this together:

ProjectRepoRole
uustreakuustreakAutomated WCAG auditing at scale
wcag-skillwcag-skillAI agent skill for detecting & fixing violations

Part 1: uustreak β€” WCAG Leaderboard for Norway

Background

Launched at ODIN 2025 (Norwegian Agency for Public Management and eGovernment). uustreak continuously audits 400+ Norwegian public sector websites and publishes a public leaderboard.

Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Scheduler  │────▢│  Playwright │────▢│   axe-core  β”‚
β”‚  (cron)     β”‚     β”‚  (Chromium) β”‚     β”‚  (WCAG 2.1) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚                    β”‚
                           β–Ό                    β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  Screenshotsβ”‚     β”‚  Violations β”‚
                    β”‚  + HTML     β”‚     β”‚  + Impact   β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚                    β”‚
                           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β–Ό
                            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                            β”‚  PostgreSQL     β”‚
                            β”‚  + GitHub Pages β”‚
                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

GitHub Actions Workflow

# .github/workflows/audit.yml
name: WCAG Audit
on:
  schedule:
    - cron: '0 2 * * 1'  # Weekly Monday 02:00 UTC
  workflow_dispatch:

jobs:
  audit:
    runs-on: ubuntu-latest
    timeout-minutes: 60
    steps:
      - uses: actions/checkout@v4
      - name: Setup Node
        uses: actions/setup-node@v4
        with:
          node-version: '20'
          cache: 'npm'
      - name: Install deps
        run: npm ci
      - name: Install Playwright
        run: npx playwright install --with-deps chromium
      - name: Run audit
        env:
          SUPABASE_URL: $
          SUPABASE_KEY: $
        run: node scripts/audit-all.js
      - name: Deploy to Pages
        if: success()
        uses: peaceiris/actions-gh-pages@v3
        with:
          github_token: $
          publish_dir: ./public

Impact-Weighted Scoring

Not all violations are equal. uustreak calculates a 0-100 score:

score = 100 - Ξ£(impactWeight Γ— nodeCount)
impactWeight: critical=10, serious=5, moderate=2, minor=0.5

This produces a fair leaderboard β€” a site with 1 critical error ranks lower than one with 50 minor errors.

Results (First 6 Months)

MetricValue
Websites audited427
Total violations found12,847
Critical violations1,203
Orgs improving score67%
Avg score improvement+12.3 points

Part 2: wcag-skill β€” AI Agent for WCAG Fixes

The Problem

Accessibility audits are slow, manual, and repetitive. Developers fix the same violation patterns (missing alt text, low contrast, improper heading order) across projects. CI catches syntax errors β€” why not accessibility errors?

Solution: wcag-skill for Hermes Agent

wcag-skill is a Hermes Agent skill that:

  1. Scans any URL or local HTML with axe-core (Playwright)
  2. Analyzes violations with LLM reasoning (context-aware fixes)
  3. Proposes concrete code patches (unified diffs)
  4. Integrates into CI/CD as a wcag-check step

Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Input      │────▢│  Scanner    │────▢│  Analyzer   β”‚
β”‚  (URL/File) β”‚     β”‚  (axe-core) β”‚     β”‚  (LLM)      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚                    β”‚
                           β–Ό                    β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  Violations β”‚     β”‚  Fix Plan   β”‚
                    β”‚  (JSON)     β”‚     β”‚  (Diff)     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚                    β”‚
                           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                     β–Ό
                            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                            β”‚  Output:        β”‚
                            β”‚  - SARIF        β”‚
                            β”‚  - PR comments  β”‚
                            β”‚  - Local patchesβ”‚
                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

LLM-Powered Fix Generation

The skill doesn't just report violations β€” it proposes fixes:

# skills/wcag_skill/analyzer.py
class WCAGAnalyzer:
    def __init__(self, llm_client):
        self.llm = llm_client
    
    async def propose_fix(self, violation: Violation, context: str) -> FixProposal:
        prompt = f"""
        WCAG Violation: {violation.rule_id} ({violation.impact})
        Element: {violation.target}
        HTML: {violation.html}
        Surrounding Context: {context}
        
        Provide a minimal fix as a unified diff. Only change what's necessary.
        """
        
        response = await self.llm.complete(prompt, model="gpt-4o-mini")
        return parse_diff(response)

Example output for missing alt text:

--- a/src/components/Hero.astro
+++ b/src/components/Hero.astro
@@ -12,7 +12,7 @@
   <div class="hero-content">
-    <img src="/hero-bg.jpg" class="hero-image" />
+    <img src="/hero-bg.jpg" class="hero-image" alt="Sunrise over Oslo fjord" />
     <h1>Welcome to Our Service</h1>
   </div>

CI/CD Integration (GitHub Actions)

# .github/workflows/wcag.yml
name: WCAG Check
on: [pull_request]

jobs:
  wcag:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Build
        run: npm run build
      - name: WCAG Skill
        uses: turbolego/wcag-skill-action@v1
        with:
          path: ./dist
          fail-on: "critical,serious"
          comment-pr: true
        env:
          HERMES_API_KEY: $

Supported Rules (WCAG 2.1 AA + AAA)

CategoryRulesAuto-fixable
Perceivablecolor-contrast, image-alt, text-spacingβœ… Most
Operablekeyboard, focus-order, focus-visible⚠️ Partial
Understandablelabel, heading-order, languageβœ… Most
Robustparse, name-role-value⚠️ Partial

Together: Audit β†’ Fix β†’ Prevent

uustreak (weekly audit) β†’ wcag-skill (PR fixes) β†’ Clean CI gate
  1. uustreak runs weekly, finds regressions
  2. wcag-skill runs on every PR, blocks new violations
  3. Developers get instant fix proposals in PR comments
  4. Leadership sees trends on public leaderboard


Accessibility is not a feature β€” it's a requirement. Automation makes it enforceable.