Skip to main content

broken.conf

nginx-broken.conf is a comprehensive nginx map configuration containing 28,822 redirect rules specifically designed to handle broken URLs, malicious crawler requests, typos, and.

Pathnginx/redirect/broken.conf
LanguageNginx
Lines29352
Sourcenginx/redirect/broken.conf
at 406d3e423

Read this when Use this page when tracing server routing, redirects, deployment configuration, or Nginx behavior around broken.

Overview

broken.conf is a comprehensive nginx map configuration containing roughly 29,600 redirect rules specifically designed to handle broken URLs, malicious crawler requests, typos, and malformed links. Unlike move.conf which handles legitimate content moves, this file serves as a defensive layer that cleans up garbage URLs and redirects broken requests to appropriate destinations.

This file represents years of accumulated fixes for broken external links, typosquatting, malicious bots, and URL corruption. It implements a "zero error log" philosophy where even garbage requests are handled gracefully rather than generating 404 errors.

File Structure

Location: nginx/redirect/broken.conf

Type: Nginx map configuration

Size: ~29,700 lines, ~1.9MB

Redirect count: ~29,600 rules

Ordering: The file header declares it is "ordered by priority, because first match wins": removed features first, then garbage-suffix regexp substitutions, then the huge "literal matches" section, then the generic "GENERAL blacklist" anti-bot cleanups at the end.

Format: Same nginx map syntax as move.conf; suppression targets carry a debug fragment tag

"~^/broken/url$" "/correct/url";
"~^/malicious/.*$" "/404#06324";

Major Categories

1. Literal Matches (bulk of the file, after the top-of-file blocks)

Purpose: Fix common broken URLs and typos

"~^//$" "/";  # Double slash → home
"~^/image/$" "/doc/index"; # Old image directory
"~^/doc/genetics/heritable/2018-prasad\.pdf.*$" "https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5994200/";

Patterns:

  • Malformed URLs (double slashes, missing paths)
  • Redirects to external sources when local copy unavailable
  • Common misspellings and variations

2. Newsletter Date Normalization (Throughout)

Extensive redirects for newsletter date variations:

"~^/newsletter/2020/00$" "/newsletter/2020/01";  # Month 00 → 01
"~^/newsletter/2020/010.*$" "/newsletter/2021/10"; # Leading zero typos
"~^/newsletter/2018/1$" "/newsletter/2018/01"; # Single digit → zero-padded

Year/month shorthand redirects (active):

"~^/2020/01$" "/newsletter/2020/01";
"~^/2021/10$" "/newsletter/2021/10";

These live entries map bare /YYYY/MM paths onto the corresponding newsletter pages, including future-dated ones through at least /2027/10.

3. Danbooru Dataset URL Variations (Lines 27-74)

Background: gwern.net hosts the Danbooru2020/2021 anime image dataset

Massive effort to catch all typo variations:

"~^/Danbooru2020.*$" "/danbooru2021#danbooru2020";
"~^/danbooru2020$" "/danbooru2021#danbooru2020";
"~^/Dabooru2020$" "/danbooru2021#danbooru2020"; # Typo
"~^/Dabbooru2020.*" "/danbooru2021#danbooru2020"; # Double-b typo
"~^/Danboru2020$" "/danbooru#danbooru2020"; # Missing 'o'
"~^/Danbooru20120$" "/danbooru2021#danbooru2020"; # Year typo
"~^/danbooru204$" "/danbooru2021#danbooru2020"; # Truncated
"~^/Danboo$" "/danbooru201"; # Severely truncated
"~^/banboor.*$" "/danbooru201"; # Wrong first letter

Future-dataset-name redirects:

"~^/[Dd]anbooru2022.*$" "/danbooru2021";
"~^/[Dd]anbooru2023.*$" "/danbooru2021";

These handle people guessing future dataset names (both point at /danbooru2021, the last released version).

4. File Format Redirects (Throughout)

Fix incorrect file extensions and formats:

"~^/path/file.png.*$" "/path/file.jpg";  # Wrong extension
"~^/path/file.jpeg.*$" "/path/file.jpg"; # JPEG → JPG normalization

Similar to nginx.conf but for more obscure/broken variants.

When local PDFs aren't available, redirect to authoritative sources:

"~^/doc/genetics/heritable/2018-prasad\.pdf.*$"
"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5994200/";

"~^/doc/psychology/1943-maslow.pdf.*$"
"https://psycnet.apa.org/journals/rev/50/4/370/";

Pattern: Local path → PubMed Central, PsycNET, Archive.org, etc.

6. Author Name Variations (Lines 10000-10026)

Extensive disambiguation for common surnames:

# Schwartz variations
"~^/1994-schwartz\.pdf.*$" "/doc/iq/1994-schwartz.pdf";
"~^/2017-schwartz\.pdf.*$" "/doc/genetics/heritable/correlation/2017-schwartz.pdf";
"~^/1997-schwartz\.pdf.*$" "/doc/statistics/bias/1997-schwartz.pdf";
"~^/1998-swartz\.pdf.*$" "/doc/fiction/1998-swartz.pdf"; # Different author

# Horowitz/Horwitz variations
"~^/2019-horowitz\.pdf.*$" "/doc/sociology/2019-horowitz.pdf";
"~^/1990-horwitz\.pdf.*$" "/doc/statistics/causality/1990-horwitz.pdf";
"~^/2018-horwitz\.pdf.*$" "/doc/genetics/heritable/2018-horwitz.pdf";

Problem: Year-author citations are ambiguous when multiple papers exist

Solution: Redirect based on most likely context or most popular paper

7. Malicious/Broken Crawler Requests ("GENERAL blacklist", second half of file)

Purpose: Prevent error log spam from broken bots. This section begins at the ## GENERAL blacklist due to broken crawlers header (~line 23,000) and runs to the end of the file. Every 404 target carries a numeric fragment tag so the firing rule can be identified when debugging:

"~^//content.php.*$" "/404#06324";
"~^/class-php.*$" "/404#06325";
"~^//bp.php.*$" "/404#06327";
"~^/.windsurf.*$" "/404#06328"; # Editor config files
"~^/.cursor.*$" "/404#06329"; # Cursor editor
"~^/smartoptimizer.*$" "/404#06330";
"~^/OVA-UN.*$" "/404#06332";
"~^/outfits.*$" "/404#06333";
"~^//modules/.*$" "/404#06334";

Blocked patterns:

  • PHP files (gwern.net is static)
  • WordPress paths (site isn't WordPress)
  • Config files (.bashrc, .profile, etc.)
  • Credentials (.credentials.json, aws.json, stripe.*)
  • Development artifacts (fly.toml, k8s.yml, meteor.settings)
  • Editor configs (.windsurf, .cursor)

8. Security-Sensitive Blocks

Explicit blocking of credential-seeking requests:

"~^/\.*_credentials\.json$" "/404#06337";
"~^/\.bash_profile.*$" "/404#06340";
"~^/\.bashrc.*$" "/404#06341";
"~^/\.profile.*$" "/404#06343";
"~^/\.sqlite3.*$" "/404#06344";
"~^/aws\.json.*$" "/404#06345";
"~^/gcp-key.*$" "/404#06351";
"~^/stripe\..*$" "/404#06358";
"~^/stripe_.*$" "/404#06359";
"~^/~/\.aws/.*$" "/404#06360";

Attack pattern: Bots scanning for exposed credentials and config files

Defense: Explicitly return 404 (could also return 403 Forbidden)

9. Removed Features (Top of file)

## removed feature:
"~^/static/previews/.*\.png$" "/404#06531";

Link previews feature was removed; old requests get a tagged 404. This block sits at the very top of the file so it wins over any later match.

10. Regex Substitution Rules (Top of file, after removed features)

Pattern cleanup using capture groups — placed early so garbage endings are stripped before the literal matches are consulted:

"~^/(?<u>.*)\.$" "$u";  # Strip trailing dots
"~^/(?<u>.*)/trackback$" "$u"; # Remove /trackback suffix
"~^/(?<u>.*)/trackback/$" "$u"; # Remove /trackback/ suffix
"~^/(?<u>.*)\'$" "$u"; # Strip trailing apostrophes

Named capture syntax:

  • (?<u>.*) - Capture everything into variable u
  • $u - Reference captured content in destination

Purpose: Clean up malformed URLs from broken crawlers/tools

11. Special Cases and Edge Cases

# Link bibliography redirect
"~^/doc/link-bibliography.*$" "/404#06532";

# Google URL mangling
"~^https:/creativecommons.org/publicdomain/zero/1.0/$"
"https:/creativecommons.org/publicdomain/zero/1.0/";

# External site misdirections
"~^/item\?id=18675280.*$" "https://news.ycombinator.com/item?id=18675280";
"~^/pubmed/26853120.*$" "https://pubmed.ncbi.nlm.nih.gov/pubmed/26853120";

Redirect Destination Types

1. Internal Redirects

Point to correct path within gwern.net:

"~^/wrong/path$" "/correct/path";

2. External Redirects

Point to authoritative external sources:

"~^/local/file.pdf$" "https://pubmed.ncbi.nlm.nih.gov/...";

3. Explicit 404s

Send obviously-malicious requests to 404, always with a fragment tag identifying the rule (the policy documented in gwern.net.conf requires every redirect-to-404 to be tagged):

"~^/malicious/pattern.*$" "/404#00016";

4. Anchor-Based Redirects

Redirect to section of a page:

"~^/old-page$" "/new-page#section";
"~^/Danbooru2000$" "/danbooru2021#danbooru2020";

URL Corruption Patterns Handled

1. Encoding Issues

"~^/Danbooru2020%E3%81.*$" "/dabooru2020";  # UTF-8 encoding garbage
"~^/danbooru2020%C3%AF%C2%BC%C5%922020$" "/danbooru2020";

Cause: Double-encoding, charset mismatches, broken link parsers

2. Path Duplication

"~^/doc/rotten.*/https/www.edge.org/conversation/...$"
"https://www.edge.org/conversation/...";

Pattern: /doc/rotten.com/ prefix incorrectly prepended to external URLs

3. Truncation

"~^/Danboo$" "/danbooru201";
"~^/banboor.*$" "/danbooru201";

Cause: Character limits in referrers, truncated bookmarks

4. Case Variations

"~^/[Dd]anbooru2022.*$" "/danbooru2021";

Pattern: Case-insensitive matching for common typos

Pattern Analysis

Common Typo Classes

Transposition:

  • dabooru instead of danbooru (transposed 'n' and 'a')

Omission:

  • danboru instead of danbooru (missing 'o')

Duplication:

  • dabbooru instead of danbooru (doubled 'b')

Truncation:

  • danbooru204 instead of danbooru2020 (incomplete year)

Newsletter Date Issues

Zero-padding confusion:

  • 2020/1 vs 2020/01
  • 2020/010 (double-digit typo)
  • 2020/00 (zero month)

Pattern: Users forget zero-padding, automation adds extra zeros

Performance Characteristics

Map Size Impact

~29,600 rules in a single nginx map:

Lookup performance:

  • Literal matches: O(1) hash lookup (fast)
  • Regex matches: O(n) sequential evaluation (slower)
  • This file is almost entirely regex, so worst-case performance

Memory footprint:

  • ~2MB config file
  • Loaded into nginx memory on startup
  • Shared across all worker processes

Mitigation strategies:

  • Most requests hit cache/CDN (redirects infrequent)
  • Regex patterns anchored (^/$) for early rejection
  • File is explicitly ordered by priority since first match wins (removed features → garbage-suffix substitutions → literal matches → generic anti-bot cleanups)

Maintenance Philosophy

Zero Error Log Policy

Goal: Every URL request should have a defined response (even if 404)

Evidence:

  • Exhaustive typo coverage
  • Malicious pattern blocking
  • Explicit 404s for removed features

Benefit: Error logs contain only genuine issues, not noise

Defensive Programming

Approach: Anticipate all possible broken inputs

Patterns:

  • Multiple redirects for same destination (typo variants)
  • Year-by-year newsletter redirects (dates 2014-2027+)
  • Security blocks for common vulnerability scanners

Historical Accumulation

Growth pattern:

  • File grows as new broken links are discovered
  • Each external citation creates potential for future typos
  • Bot attacks add new malicious patterns to block

Evidence: Temporal comments about removing redirects in future years

Integration with gwern.net.conf

Division of Responsibility

FilePurposeExample
move.confLegitimate content movesReorganized files, renamed pages
broken.confBroken/malicious URLsTypos, bots, malformed requests

Combined Processing

Both files are included in the same nginx map in gwern.net.conf:

map $request_uri $new_uri {
include /home/gwern/gwern.net/static/nginx/redirect/move.conf;
include /home/gwern/gwern.net/static/nginx/redirect/broken.conf;
}

Evaluation order:

  1. First match wins
  2. move.conf rules are evaluated first (legitimate redirects)
  3. broken.conf catches remaining broken patterns
  4. Default (empty) means no redirect → serve normally or 404

Security Considerations

Bot Protection

Blocks common vulnerability scanner patterns:

  • WordPress paths (/wp-admin, /modules/)
  • PHP files (site is static)
  • Config files (credentials, env files)
  • Infrastructure configs (Kubernetes, Docker, etc.)

Goal: Reduce attack surface and log noise

Credential Harvesting Defense

Explicit blocks for:

  • AWS credentials
  • GCP keys
  • Stripe API keys
  • Generic _credentials.json
  • Shell config files

Attack vector: Automated scanners looking for exposed secrets

Rate Limit Considerations

Impact: Each redirect is a server response

Mitigation:

  • Rate limiting on IP level (separate nginx config)
  • CDN/cache layer prevents most redirect hits
  • 404 responses are cheap (no disk I/O)

Temporal Redirects

Future-Dated Patterns

"~^/[Dd]anbooru2022.*$" "/danbooru2021";
"~^/[Dd]anbooru2023.*$" "/danbooru2021";

Strategy: Catch people guessing future dataset names (all mapped to the last released version)

Alternative approach: Generic pattern like "~^/[Dd]anbooru20[2-9][0-9].*$" (but less precise)

Newsletter Redirects for Future Years

"~^/2025/10$" "/newsletter/2025/10";
"~^/2026/10$" "/newsletter/2026/10";
"~^/2027/10$" "/newsletter/2027/10";

Status: Active entries

Reason: Pre-empt the recurring /YYYY/MM shorthand mistake for future newsletter issues

Common Failure Modes Addressed

1. Copy-Paste Errors

Users copy URL with trailing punctuation:

"~^/(?<u>.*)\'$" "$u";  # Strip trailing apostrophe
"~^/(?<u>.*)\.$" "$u"; # Strip trailing period

2. Referrer Truncation

Referrer headers or bookmarks truncate URLs:

"~^/Danboo$" "/danbooru201";

3. URL Encoding Corruption

Double-encoding or charset issues:

"~^/Danbooru2020%E3%81.*$" "/dabooru2020";

4. Autocorrect Interference

User's device autocorrects URL:

"~^/Danbooru2020It$" "/danbooru2021#danbooru2020";  # "It" autocorrect

Regex Patterns Used

Named Captures

"~^/(?<u>.*)\.$" "$u";

Syntax: (?<name>pattern) captures into $name

Wildcards and Anchors

"~^/exact-prefix.*$" "/destination";
  • ^ - Start of URL
  • .* - Any characters
  • $ - End of URL

Character Classes

"~^/[Dd]anbooru2022.*$" "/danbooru2021";
  • [Dd] - Matches 'D' or 'd'

Escape Sequences

"~^/\.*_credentials\.json$" "/404#06337";
  • \. - Literal dot
  • \.* - Any number of literal dots (e.g., ._credentials.json)

Maintenance Workflow

Adding New Broken URL Fixes

Process:

  1. Monitor 404 hits (gwern.net exposes them at /metadata/404-hits.txt.log, fed to a redirectGuesser tool)
  2. Identify patterns in broken URLs
  3. Add redirect rule to the appropriate section of broken.conf
  4. Test redirect
  5. Deploy and monitor

Example workflow:

# Find common 404 patterns
curl --silent "https://gwern.net/metadata/404-hits.txt.log" | sort | uniq -c | sort -rn

# Add redirect rule (in the right section; tag /404 targets)
echo '"~^/broken-pattern.*$" "/correct-path";' >> broken.conf

# Test configuration
nginx -t

# Reload
nginx -s reload

Cleanup Strategy

When to remove rules:

  • Temporal redirects past their expiration date
  • Redirects with zero hits for extended period
  • Superseded by broader pattern matches

Challenges:

  • Hard to know if redirect still needed
  • External sites may link to old URLs indefinitely
  • Conservative approach: keep everything

Statistics and Insights

MetricValue
Total redirects~29,600
File size~1.9 MB
Lines~29,700
Literal 404 blocks~50+ security-related
Danbooru typo variants15+ variations
Newsletter date fixes100+ date variations
External redirectsHundreds (to PubMed, LibGen, etc.)
Author name disambiguationsDozens

Pattern distribution (estimated):

  • 40% File path corrections and format fixes
  • 25% Typo and variation handling
  • 20% External source redirects
  • 10% Security blocks (bots, malicious requests)
  • 5% URL corruption cleanup

See Also

Philosophy: Embrace the Chaos

This file represents a pragmatic approach to the messy reality of the web:

Accept that URLs will break:

  • Typos happen
  • Bots will misbehave
  • URLs get corrupted in transit
  • Users copy-paste carelessly

Handle it gracefully:

  • Redirect to correct content when possible
  • Return clean 404s for malicious requests
  • Reduce error log noise
  • Maintain link graph integrity

Learn from history:

  • Each broken URL teaches a pattern
  • Accumulate fixes over time
  • Build comprehensive coverage

Result: A robust, resilient link structure that degrades gracefully even when users (or bots) provide garbage input.