feat(curation): data quality pipeline — Phases 1-3

Add comprehensive data curation system to clean up the 197K skill
dataset and show only quality browse-ready skills to users.

Phase 1 — Database exploration:
- Explore scripts (explore.ts, explore.mjs, explore.sql) for analysis
- Discovered: 69% duplicates, 77% aggregator/fork noise

Phase 2 — Data cleanup and classification:
- Schema: 6 new curation columns + 4 indexes
- curate.mjs: 8-step pipeline (classify, dedup, fork detection, etc.)
- Result: 197K → 60K unique → 16K browse-ready skills
- Bug fix: securityStatus was computed but never stored during crawl

Phase 3 — UI browse-ready filters:
- browseReadyFilter applied to 17+ query functions
- Homepage stats show accurate browse-ready counts
- Stats API filtered (previously had no WHERE clause)
- Category counts recalculated (e.g. 45K → 3.1K)
- Featured skills exclude duplicates and aggregators

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
airano
2026-02-19 14:31:25 +03:30
parent 31f5df0900
commit caca09fbe7
12 changed files with 2615 additions and 28 deletions

View File

@@ -67,6 +67,19 @@ export const skills = pgTable(
isBlocked: boolean('is_blocked').default(false), // Blocked from re-indexing (owner requested removal)
lastScanned: timestamp('last_scanned'),
// Curation (populated by batch scripts, not crawler)
qualityScore: integer('quality_score'), // 0-100 from analyzer
qualityDetails: jsonb('quality_details').$type<{
documentation: number;
maintenance: number;
popularity: number;
factors: Array<{ name: string; score: number; weight: number; details?: string }>;
}>(),
skillType: text('skill_type').$type<'standalone' | 'project-bound' | 'collection' | 'aggregator'>(),
isDuplicate: boolean('is_duplicate').default(false),
canonicalSkillId: text('canonical_skill_id'), // points to the "original" if this is a duplicate
repoSkillCount: integer('repo_skill_count'), // cached count of skills in same repo
// Content (cached)
contentHash: text('content_hash'),
rawContent: text('raw_content'),
@@ -103,6 +116,10 @@ export const skills = pgTable(
updatedIdx: index('idx_skills_updated').on(table.updatedAt),
sourceFormatIdx: index('idx_skills_source_format').on(table.sourceFormat),
lastDownloadedIdx: index('idx_skills_last_downloaded').on(table.lastDownloadedAt),
qualityIdx: index('idx_skills_quality').on(table.qualityScore),
skillTypeIdx: index('idx_skills_type').on(table.skillType),
duplicateIdx: index('idx_skills_duplicate').on(table.isDuplicate),
contentHashIdx: index('idx_skills_content_hash').on(table.contentHash),
})
);