Add automated article processing system
🤖 Automated workflow that: - Monitors aiadopters.club RSS feed every 6 hours - Extracts claim data using Claude AI - Creates PRs for review before going live Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
parent
e127f885ec
commit
b1ef96c48b
|
|
@ -0,0 +1,99 @@
|
||||||
|
name: Auto-Add New Articles
|
||||||
|
|
||||||
|
on:
|
||||||
|
# Run every 6 hours
|
||||||
|
schedule:
|
||||||
|
- cron: '0 */6 * * *'
|
||||||
|
|
||||||
|
# Allow manual trigger
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
check-and-add-articles:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
|
||||||
|
permissions:
|
||||||
|
contents: write
|
||||||
|
pull-requests: write
|
||||||
|
|
||||||
|
steps:
|
||||||
|
- name: Checkout repository
|
||||||
|
uses: actions/checkout@v4
|
||||||
|
|
||||||
|
- name: Setup Node.js
|
||||||
|
uses: actions/setup-node@v4
|
||||||
|
with:
|
||||||
|
node-version: '20'
|
||||||
|
cache: 'npm'
|
||||||
|
|
||||||
|
- name: Install dependencies
|
||||||
|
run: npm ci
|
||||||
|
|
||||||
|
- name: Check for new articles and process
|
||||||
|
env:
|
||||||
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
||||||
|
run: |
|
||||||
|
echo "🔍 Checking for new articles..."
|
||||||
|
npx tsx scripts/auto-add-new-articles.ts
|
||||||
|
|
||||||
|
- name: Check for changes
|
||||||
|
id: check_changes
|
||||||
|
run: |
|
||||||
|
if git diff --quiet src/data/claims.ts; then
|
||||||
|
echo "has_changes=false" >> $GITHUB_OUTPUT
|
||||||
|
echo "No changes detected"
|
||||||
|
else
|
||||||
|
echo "has_changes=true" >> $GITHUB_OUTPUT
|
||||||
|
echo "Changes detected in claims.ts"
|
||||||
|
fi
|
||||||
|
|
||||||
|
- name: Build site
|
||||||
|
if: steps.check_changes.outputs.has_changes == 'true'
|
||||||
|
run: npm run build
|
||||||
|
|
||||||
|
- name: Create Pull Request
|
||||||
|
if: steps.check_changes.outputs.has_changes == 'true'
|
||||||
|
uses: peter-evans/create-pull-request@v6
|
||||||
|
with:
|
||||||
|
token: ${{ secrets.GITHUB_TOKEN }}
|
||||||
|
commit-message: 'Add new article(s) from aiadopters.club'
|
||||||
|
branch: auto-add-articles
|
||||||
|
delete-branch: true
|
||||||
|
title: '🤖 Auto-add new article(s) from aiadopters.club'
|
||||||
|
body: |
|
||||||
|
## 🆕 New Article(s) Detected
|
||||||
|
|
||||||
|
This PR was automatically created by the article monitoring workflow.
|
||||||
|
|
||||||
|
### What Changed
|
||||||
|
- ✅ New article(s) added to `src/data/claims.ts`
|
||||||
|
- ✅ Build completed successfully
|
||||||
|
|
||||||
|
### Review Checklist
|
||||||
|
- [ ] Review extracted claims for accuracy
|
||||||
|
- [ ] Check claim titles are concise and clear
|
||||||
|
- [ ] Verify topics are appropriate
|
||||||
|
- [ ] Confirm statistics have proper context
|
||||||
|
- [ ] Review supporting context paragraph
|
||||||
|
|
||||||
|
### Next Steps
|
||||||
|
1. Review the changes in `src/data/claims.ts`
|
||||||
|
2. If everything looks good, merge this PR
|
||||||
|
3. Deploy: `netlify deploy --dir=out --prod`
|
||||||
|
|
||||||
|
---
|
||||||
|
🤖 Generated by GitHub Actions
|
||||||
|
labels: |
|
||||||
|
automated
|
||||||
|
content
|
||||||
|
|
||||||
|
- name: Summary
|
||||||
|
if: steps.check_changes.outputs.has_changes == 'true'
|
||||||
|
run: |
|
||||||
|
echo "✅ New article(s) processed and PR created!"
|
||||||
|
echo "📝 Review the PR and merge when ready"
|
||||||
|
|
||||||
|
- name: No changes summary
|
||||||
|
if: steps.check_changes.outputs.has_changes == 'false'
|
||||||
|
run: |
|
||||||
|
echo "✨ No new articles found - database is up to date!"
|
||||||
|
|
@ -0,0 +1,206 @@
|
||||||
|
# Automated Article Processing Setup
|
||||||
|
|
||||||
|
This system automatically monitors aiadopters.club for new articles and creates claim pages automatically.
|
||||||
|
|
||||||
|
## 🎯 How It Works
|
||||||
|
|
||||||
|
1. **GitHub Actions** runs every 6 hours (or manually triggered)
|
||||||
|
2. **Checks RSS feed** for new articles not in your database
|
||||||
|
3. **Extracts data** using Claude AI (claims, statistics, quotes, etc.)
|
||||||
|
4. **Adds to claims.ts** automatically
|
||||||
|
5. **Creates Pull Request** for you to review
|
||||||
|
6. **You review & merge** when ready
|
||||||
|
|
||||||
|
## 🔧 Setup Instructions
|
||||||
|
|
||||||
|
### Step 1: Add Anthropic API Key to GitHub Secrets
|
||||||
|
|
||||||
|
1. Get your Anthropic API key:
|
||||||
|
- Go to https://console.anthropic.com/
|
||||||
|
- Sign in or create account
|
||||||
|
- Navigate to "API Keys"
|
||||||
|
- Create a new key (keep it safe!)
|
||||||
|
|
||||||
|
2. Add to GitHub repository:
|
||||||
|
- Go to your repository: https://github.com/kbanc85/kbanc-nextjs
|
||||||
|
- Click "Settings" tab
|
||||||
|
- In left sidebar, click "Secrets and variables" → "Actions"
|
||||||
|
- Click "New repository secret"
|
||||||
|
- Name: `ANTHROPIC_API_KEY`
|
||||||
|
- Value: (paste your API key)
|
||||||
|
- Click "Add secret"
|
||||||
|
|
||||||
|
### Step 2: Push the Workflow Files
|
||||||
|
|
||||||
|
The automation files are already created. Push them to GitHub:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add .
|
||||||
|
git commit -m "Add automated article processing workflow"
|
||||||
|
git push
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 3: Enable GitHub Actions
|
||||||
|
|
||||||
|
1. Go to your repository on GitHub
|
||||||
|
2. Click "Actions" tab
|
||||||
|
3. If prompted, enable workflows
|
||||||
|
|
||||||
|
### Step 4: Test the Automation (Optional)
|
||||||
|
|
||||||
|
Manual trigger to test immediately:
|
||||||
|
|
||||||
|
1. Go to "Actions" tab
|
||||||
|
2. Click "Auto-Add New Articles" workflow
|
||||||
|
3. Click "Run workflow" button
|
||||||
|
4. Select branch (main/master)
|
||||||
|
5. Click green "Run workflow" button
|
||||||
|
|
||||||
|
The workflow will:
|
||||||
|
- Check RSS feed
|
||||||
|
- Process any new articles
|
||||||
|
- Create a PR if articles found
|
||||||
|
- Or report "no new articles"
|
||||||
|
|
||||||
|
## 📅 Schedule
|
||||||
|
|
||||||
|
**Automatic runs:** Every 6 hours
|
||||||
|
- 12:00 AM UTC
|
||||||
|
- 6:00 AM UTC
|
||||||
|
- 12:00 PM UTC
|
||||||
|
- 6:00 PM UTC
|
||||||
|
|
||||||
|
**Manual runs:** Anytime via GitHub Actions UI
|
||||||
|
|
||||||
|
## 🔍 What Gets Automated
|
||||||
|
|
||||||
|
### Automatically Extracted:
|
||||||
|
- ✅ Article title and date
|
||||||
|
- ✅ URL-friendly slug
|
||||||
|
- ✅ Featured claim (homepage summary)
|
||||||
|
- ✅ Description for grid view
|
||||||
|
- ✅ 3-4 key points
|
||||||
|
- ✅ 5 atomic claims with short titles
|
||||||
|
- ✅ Topic categorization
|
||||||
|
- ✅ Key quote with attribution
|
||||||
|
- ✅ 2-4 statistics with context
|
||||||
|
- ✅ Supporting context paragraph
|
||||||
|
|
||||||
|
### Your Review:
|
||||||
|
The workflow creates a **Pull Request** for you to review before merging.
|
||||||
|
|
||||||
|
**Why a PR?** So you can:
|
||||||
|
- ✅ Verify extracted data is accurate
|
||||||
|
- ✅ Edit claims if needed
|
||||||
|
- ✅ Check topic categorization
|
||||||
|
- ✅ Ensure quality standards
|
||||||
|
- ✅ Merge when satisfied
|
||||||
|
|
||||||
|
## 📝 Manual Processing (If Needed)
|
||||||
|
|
||||||
|
You can still manually process articles:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Process a single article
|
||||||
|
npx tsx scripts/extract-article-data.ts https://aiadopters.club/p/article-slug
|
||||||
|
|
||||||
|
# Check for new articles only
|
||||||
|
npx tsx scripts/check-new-articles.ts
|
||||||
|
|
||||||
|
# Run full automation
|
||||||
|
npx tsx scripts/auto-add-new-articles.ts
|
||||||
|
```
|
||||||
|
|
||||||
|
## 🛠 Workflow Files
|
||||||
|
|
||||||
|
- `.github/workflows/auto-add-articles.yml` - GitHub Actions workflow
|
||||||
|
- `scripts/check-new-articles.ts` - RSS feed monitoring
|
||||||
|
- `scripts/extract-article-data.ts` - AI extraction logic
|
||||||
|
- `scripts/auto-add-new-articles.ts` - Main orchestration
|
||||||
|
- `scripts/add-claim-to-data.ts` - Add to claims.ts (already existed)
|
||||||
|
|
||||||
|
## 🔔 Notifications
|
||||||
|
|
||||||
|
You'll get notifications when:
|
||||||
|
- ✅ New article PRs are created
|
||||||
|
- ❌ Workflow fails (API issues, parsing errors, etc.)
|
||||||
|
|
||||||
|
Configure in GitHub: Settings → Notifications → Actions
|
||||||
|
|
||||||
|
## 💰 Cost Estimate
|
||||||
|
|
||||||
|
**Anthropic API pricing:**
|
||||||
|
- ~2,000-4,000 tokens per article (input)
|
||||||
|
- ~1,000 tokens per article (output)
|
||||||
|
- With Claude Sonnet: ~$0.05-0.10 per article
|
||||||
|
|
||||||
|
**Expected monthly cost:**
|
||||||
|
- If 2 articles/week: ~$0.40-0.80/month
|
||||||
|
- If 1 article/day: ~$1.50-3.00/month
|
||||||
|
|
||||||
|
**GitHub Actions:** Free (within generous limits)
|
||||||
|
|
||||||
|
## 🚨 Troubleshooting
|
||||||
|
|
||||||
|
### Workflow doesn't run
|
||||||
|
- Check GitHub Actions is enabled
|
||||||
|
- Verify ANTHROPIC_API_KEY secret exists
|
||||||
|
- Check workflow file is in `.github/workflows/`
|
||||||
|
|
||||||
|
### Extraction fails
|
||||||
|
- Check API key is valid
|
||||||
|
- Look at workflow logs for error details
|
||||||
|
- Article might be behind paywall
|
||||||
|
- API might be down temporarily
|
||||||
|
|
||||||
|
### No PR created
|
||||||
|
- No new articles found (check RSS feed)
|
||||||
|
- Changes failed to save
|
||||||
|
- Check workflow logs for details
|
||||||
|
|
||||||
|
### Build fails
|
||||||
|
- TypeScript errors in extracted data
|
||||||
|
- Run locally: `npm run build`
|
||||||
|
- Check claims.ts formatting
|
||||||
|
|
||||||
|
## 🎓 How to Modify
|
||||||
|
|
||||||
|
### Change schedule
|
||||||
|
Edit `.github/workflows/auto-add-articles.yml`:
|
||||||
|
```yaml
|
||||||
|
schedule:
|
||||||
|
- cron: '0 */6 * * *' # Every 6 hours
|
||||||
|
# Examples:
|
||||||
|
# - cron: '0 */3 * * *' # Every 3 hours
|
||||||
|
# - cron: '0 0 * * *' # Daily at midnight
|
||||||
|
# - cron: '0 9 * * 1' # Every Monday at 9 AM
|
||||||
|
```
|
||||||
|
|
||||||
|
### Adjust extraction prompt
|
||||||
|
Edit `scripts/extract-article-data.ts` → `callAnthropicAPI()` function
|
||||||
|
|
||||||
|
### Change AI model
|
||||||
|
Edit `scripts/extract-article-data.ts`:
|
||||||
|
```typescript
|
||||||
|
model: 'claude-sonnet-4-5-20250929' // Current
|
||||||
|
// or
|
||||||
|
model: 'claude-sonnet-3-5-20241022' // Older, cheaper
|
||||||
|
```
|
||||||
|
|
||||||
|
## ✅ Verification
|
||||||
|
|
||||||
|
After setup, verify:
|
||||||
|
- [ ] ANTHROPIC_API_KEY added to GitHub Secrets
|
||||||
|
- [ ] Workflow files pushed to repository
|
||||||
|
- [ ] GitHub Actions enabled
|
||||||
|
- [ ] Manual test run successful (optional)
|
||||||
|
|
||||||
|
## 🎉 You're Done!
|
||||||
|
|
||||||
|
The system will now automatically:
|
||||||
|
1. Monitor aiadopters.club every 6 hours
|
||||||
|
2. Process new articles with AI
|
||||||
|
3. Create PRs for your review
|
||||||
|
4. Wait for your approval before going live
|
||||||
|
|
||||||
|
**No more manual article processing!** Just review and merge PRs when they appear.
|
||||||
|
|
@ -8,7 +8,10 @@
|
||||||
"build": "next build",
|
"build": "next build",
|
||||||
"start": "next start",
|
"start": "next start",
|
||||||
"lint": "next lint",
|
"lint": "next lint",
|
||||||
"add-claim-from-json": "tsx scripts/add-claim-to-data.ts"
|
"add-claim-from-json": "tsx scripts/add-claim-to-data.ts",
|
||||||
|
"check-new-articles": "tsx scripts/check-new-articles.ts",
|
||||||
|
"extract-article": "tsx scripts/extract-article-data.ts",
|
||||||
|
"auto-add-articles": "tsx scripts/auto-add-new-articles.ts"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@hookform/resolvers": "^5.2.2",
|
"@hookform/resolvers": "^5.2.2",
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,63 @@
|
||||||
|
// Main automation script - checks for new articles and processes them
|
||||||
|
// This is the entry point called by GitHub Actions
|
||||||
|
|
||||||
|
import { checkNewArticles } from './check-new-articles';
|
||||||
|
import { extractAndAddArticle } from './extract-article-data';
|
||||||
|
|
||||||
|
async function main() {
|
||||||
|
console.log('🚀 Starting automated article processing...\n');
|
||||||
|
|
||||||
|
try {
|
||||||
|
// Step 1: Check for new articles
|
||||||
|
const newArticles = await checkNewArticles();
|
||||||
|
|
||||||
|
if (newArticles.length === 0) {
|
||||||
|
console.log('\n✅ All articles already processed - no action needed');
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Step 2: Process each new article
|
||||||
|
console.log(`\n🔄 Processing ${newArticles.length} new article(s)...\n`);
|
||||||
|
|
||||||
|
for (const article of newArticles) {
|
||||||
|
console.log(`\n${'='.repeat(80)}`);
|
||||||
|
console.log(`Processing: ${article.title}`);
|
||||||
|
console.log(`${'='.repeat(80)}`);
|
||||||
|
|
||||||
|
try {
|
||||||
|
await extractAndAddArticle(article.link);
|
||||||
|
console.log(`✅ Successfully processed: ${article.title}`);
|
||||||
|
} catch (error) {
|
||||||
|
console.error(`❌ Failed to process "${article.title}":`, error);
|
||||||
|
// Continue with next article even if one fails
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
console.log(`\n${'='.repeat(80)}`);
|
||||||
|
console.log('🎉 Automation complete!');
|
||||||
|
console.log(`${'='.repeat(80)}`);
|
||||||
|
console.log(`\n📊 Summary:`);
|
||||||
|
console.log(` Total new articles found: ${newArticles.length}`);
|
||||||
|
console.log(` Successfully processed: ${newArticles.length}`);
|
||||||
|
console.log(`\n💡 Next steps:`);
|
||||||
|
console.log(` 1. Review changes in claims.ts`);
|
||||||
|
console.log(` 2. Run build: npm run build`);
|
||||||
|
console.log(` 3. Commit and deploy if everything looks good`);
|
||||||
|
|
||||||
|
} catch (error) {
|
||||||
|
console.error('\n❌ Automation failed:', error);
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Run if called directly
|
||||||
|
if (require.main === module) {
|
||||||
|
main()
|
||||||
|
.then(() => process.exit(0))
|
||||||
|
.catch(error => {
|
||||||
|
console.error(error);
|
||||||
|
process.exit(1);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export { main };
|
||||||
|
|
@ -0,0 +1,98 @@
|
||||||
|
// Script to check RSS feed for new articles not in claims.ts
|
||||||
|
// Returns new article URLs that need to be processed
|
||||||
|
|
||||||
|
import https from 'https';
|
||||||
|
import { ALL_CLAIMS_DATA } from '../src/data/claims';
|
||||||
|
|
||||||
|
interface RSSItem {
|
||||||
|
title: string;
|
||||||
|
link: string;
|
||||||
|
pubDate: string;
|
||||||
|
guid: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
function fetchRSS(url: string): Promise<string> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
https.get(url, (res) => {
|
||||||
|
let data = '';
|
||||||
|
res.on('data', (chunk) => data += chunk);
|
||||||
|
res.on('end', () => resolve(data));
|
||||||
|
}).on('error', reject);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseRSS(xml: string): RSSItem[] {
|
||||||
|
const items: RSSItem[] = [];
|
||||||
|
const itemRegex = /<item>([\s\S]*?)<\/item>/g;
|
||||||
|
let match;
|
||||||
|
|
||||||
|
while ((match = itemRegex.exec(xml)) !== null) {
|
||||||
|
const itemContent = match[1];
|
||||||
|
|
||||||
|
const titleMatch = /<title><!\[CDATA\[(.*?)\]\]><\/title>/.exec(itemContent) ||
|
||||||
|
/<title>(.*?)<\/title>/.exec(itemContent);
|
||||||
|
const linkMatch = /<link>(.*?)<\/link>/.exec(itemContent);
|
||||||
|
const pubDateMatch = /<pubDate>(.*?)<\/pubDate>/.exec(itemContent);
|
||||||
|
const guidMatch = /<guid.*?>(.*?)<\/guid>/.exec(itemContent);
|
||||||
|
|
||||||
|
if (titleMatch && linkMatch) {
|
||||||
|
items.push({
|
||||||
|
title: titleMatch[1],
|
||||||
|
link: linkMatch[1].trim(),
|
||||||
|
pubDate: pubDateMatch ? pubDateMatch[1] : '',
|
||||||
|
guid: guidMatch ? guidMatch[1] : linkMatch[1].trim(),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return items;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function checkNewArticles(): Promise<RSSItem[]> {
|
||||||
|
console.log('🔍 Checking aiadopters.club RSS feed for new articles...');
|
||||||
|
|
||||||
|
// Fetch RSS feed
|
||||||
|
const rssUrl = 'https://aiadopters.club/feed';
|
||||||
|
const xmlData = await fetchRSS(rssUrl);
|
||||||
|
|
||||||
|
// Parse RSS items
|
||||||
|
const allArticles = parseRSS(xmlData);
|
||||||
|
console.log(`📰 Found ${allArticles.length} total articles in feed`);
|
||||||
|
|
||||||
|
// Get existing article URLs from claims.ts
|
||||||
|
const existingUrls = new Set(ALL_CLAIMS_DATA.map(claim => claim.originalUrl));
|
||||||
|
console.log(`✅ Already have ${existingUrls.size} articles in claims database`);
|
||||||
|
|
||||||
|
// Filter for new articles only
|
||||||
|
const newArticles = allArticles.filter(article => !existingUrls.has(article.link));
|
||||||
|
|
||||||
|
if (newArticles.length === 0) {
|
||||||
|
console.log('✨ No new articles found - database is up to date!');
|
||||||
|
} else {
|
||||||
|
console.log(`🆕 Found ${newArticles.length} new article(s):`);
|
||||||
|
newArticles.forEach(article => {
|
||||||
|
console.log(` - ${article.title}`);
|
||||||
|
console.log(` URL: ${article.link}`);
|
||||||
|
console.log(` Published: ${article.pubDate}`);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
return newArticles;
|
||||||
|
}
|
||||||
|
|
||||||
|
// CLI usage
|
||||||
|
if (require.main === module) {
|
||||||
|
checkNewArticles()
|
||||||
|
.then(newArticles => {
|
||||||
|
// Output as JSON for GitHub Actions to consume
|
||||||
|
console.log('\n📋 JSON Output:');
|
||||||
|
console.log(JSON.stringify(newArticles, null, 2));
|
||||||
|
process.exit(newArticles.length > 0 ? 1 : 0); // Exit code 1 if new articles found
|
||||||
|
})
|
||||||
|
.catch(error => {
|
||||||
|
console.error('❌ Error checking feed:', error);
|
||||||
|
process.exit(1);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export { checkNewArticles, RSSItem };
|
||||||
|
|
@ -0,0 +1,222 @@
|
||||||
|
// Script to extract claim data from an article using AI
|
||||||
|
// Uses Anthropic Claude API to analyze article and extract structured data
|
||||||
|
|
||||||
|
import https from 'https';
|
||||||
|
import { addClaimToDataFile } from './add-claim-to-data';
|
||||||
|
|
||||||
|
interface ExtractedClaimData {
|
||||||
|
slug: string;
|
||||||
|
title: string;
|
||||||
|
date: string;
|
||||||
|
featuredClaim: string;
|
||||||
|
description: string;
|
||||||
|
keyPoints: string[];
|
||||||
|
topics: string[];
|
||||||
|
claims: string[];
|
||||||
|
claimTitles: string[];
|
||||||
|
originalUrl: string;
|
||||||
|
quote: string;
|
||||||
|
keyStatistics: Array<{ stat: string; context: string }>;
|
||||||
|
supportingContext: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
function fetchArticleContent(url: string): Promise<string> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
https.get(url, (res) => {
|
||||||
|
let data = '';
|
||||||
|
res.on('data', (chunk) => data += chunk);
|
||||||
|
res.on('end', () => resolve(data));
|
||||||
|
}).on('error', reject);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function cleanHTML(html: string): string {
|
||||||
|
// Remove script and style tags
|
||||||
|
let text = html.replace(/<script\b[^<]*(?:(?!<\/script>)<[^<]*)*<\/script>/gi, '');
|
||||||
|
text = text.replace(/<style\b[^<]*(?:(?!<\/style>)<[^<]*)*<\/style>/gi, '');
|
||||||
|
|
||||||
|
// Remove HTML tags but keep content
|
||||||
|
text = text.replace(/<[^>]+>/g, ' ');
|
||||||
|
|
||||||
|
// Decode HTML entities
|
||||||
|
text = text.replace(/ /g, ' ');
|
||||||
|
text = text.replace(/&/g, '&');
|
||||||
|
text = text.replace(/</g, '<');
|
||||||
|
text = text.replace(/>/g, '>');
|
||||||
|
text = text.replace(/"/g, '"');
|
||||||
|
|
||||||
|
// Clean up whitespace
|
||||||
|
text = text.replace(/\s+/g, ' ').trim();
|
||||||
|
|
||||||
|
return text;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function callAnthropicAPI(articleContent: string, articleUrl: string): Promise<ExtractedClaimData> {
|
||||||
|
const apiKey = process.env.ANTHROPIC_API_KEY;
|
||||||
|
|
||||||
|
if (!apiKey) {
|
||||||
|
throw new Error('ANTHROPIC_API_KEY environment variable is required');
|
||||||
|
}
|
||||||
|
|
||||||
|
const prompt = `You are analyzing an article from aiadopters.club to extract structured claim data.
|
||||||
|
|
||||||
|
Article URL: ${articleUrl}
|
||||||
|
|
||||||
|
Article Content:
|
||||||
|
${articleContent.substring(0, 50000)} // Limit to prevent token overflow
|
||||||
|
|
||||||
|
Extract the following information in valid JSON format:
|
||||||
|
|
||||||
|
1. slug: Create a URL-friendly slug (lowercase, hyphens, no spaces) based on the title
|
||||||
|
2. title: The full article title
|
||||||
|
3. date: Today's date in YYYY-MM-DD format
|
||||||
|
4. featuredClaim: A one-sentence summary (max 120 chars) highlighting the key insight
|
||||||
|
5. description: A short description (2-3 sentences) for grid view
|
||||||
|
6. keyPoints: An array of 3-4 bullet points covering main takeaways
|
||||||
|
7. topics: Array of topic IDs from: STRATEGY, TOOLS, BUSINESS, IMPLEMENTATION, MEASUREMENT (choose 1-3 most relevant)
|
||||||
|
8. claims: Array of exactly 5 atomic claims that are independently verifiable with evidence from the article
|
||||||
|
9. claimTitles: Array of exactly 5 short headers (3-6 words each) for the claims (e.g., "AI accelerates existing developer expertise")
|
||||||
|
10. quote: One powerful quote from the article (with context if needed)
|
||||||
|
11. keyStatistics: Array of 2-4 key statistics, each with "stat" (the number/metric) and "context" (explanation)
|
||||||
|
12. supportingContext: A paragraph (3-5 sentences) explaining the methodology, research basis, or how practitioners can apply these insights
|
||||||
|
|
||||||
|
IMPORTANT RULES:
|
||||||
|
- claims and claimTitles arrays MUST have exactly 5 items each
|
||||||
|
- Claims should be specific, verifiable statements from the article
|
||||||
|
- Claim titles should be concise headers that summarize each claim
|
||||||
|
- featuredClaim should be compelling and highlight the most important insight
|
||||||
|
- Topics should reflect the actual content (AI strategy, tools, business applications, implementation details, or measurement/ROI)
|
||||||
|
- Statistics should include both the number and clear context
|
||||||
|
- Keep all text professional and evidence-based
|
||||||
|
|
||||||
|
Return ONLY valid JSON, no other text:`;
|
||||||
|
|
||||||
|
const requestData = JSON.stringify({
|
||||||
|
model: 'claude-sonnet-4-5-20250929',
|
||||||
|
max_tokens: 4000,
|
||||||
|
messages: [{
|
||||||
|
role: 'user',
|
||||||
|
content: prompt
|
||||||
|
}]
|
||||||
|
});
|
||||||
|
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const options = {
|
||||||
|
hostname: 'api.anthropic.com',
|
||||||
|
path: '/v1/messages',
|
||||||
|
method: 'POST',
|
||||||
|
headers: {
|
||||||
|
'Content-Type': 'application/json',
|
||||||
|
'x-api-key': apiKey,
|
||||||
|
'anthropic-version': '2023-06-01',
|
||||||
|
'Content-Length': Buffer.byteLength(requestData)
|
||||||
|
}
|
||||||
|
};
|
||||||
|
|
||||||
|
const req = https.request(options, (res) => {
|
||||||
|
let data = '';
|
||||||
|
res.on('data', (chunk) => data += chunk);
|
||||||
|
res.on('end', () => {
|
||||||
|
try {
|
||||||
|
const response = JSON.parse(data);
|
||||||
|
|
||||||
|
if (response.error) {
|
||||||
|
reject(new Error(`Anthropic API error: ${response.error.message}`));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!response.content || !response.content[0] || !response.content[0].text) {
|
||||||
|
reject(new Error('Unexpected API response format'));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const responseText = response.content[0].text;
|
||||||
|
|
||||||
|
// Extract JSON from response (in case there's extra text)
|
||||||
|
const jsonMatch = responseText.match(/\{[\s\S]*\}/);
|
||||||
|
if (!jsonMatch) {
|
||||||
|
reject(new Error('No JSON found in API response'));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const extractedData: ExtractedClaimData = JSON.parse(jsonMatch[0]);
|
||||||
|
|
||||||
|
// Add originalUrl
|
||||||
|
extractedData.originalUrl = articleUrl;
|
||||||
|
|
||||||
|
// Validate required fields
|
||||||
|
if (!extractedData.claims || extractedData.claims.length !== 5) {
|
||||||
|
reject(new Error(`Expected 5 claims, got ${extractedData.claims?.length || 0}`));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!extractedData.claimTitles || extractedData.claimTitles.length !== 5) {
|
||||||
|
reject(new Error(`Expected 5 claim titles, got ${extractedData.claimTitles?.length || 0}`));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
resolve(extractedData);
|
||||||
|
} catch (error) {
|
||||||
|
reject(new Error(`Failed to parse API response: ${error}`));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
req.on('error', reject);
|
||||||
|
req.write(requestData);
|
||||||
|
req.end();
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function extractAndAddArticle(articleUrl: string): Promise<void> {
|
||||||
|
console.log(`\n🤖 Processing article: ${articleUrl}`);
|
||||||
|
|
||||||
|
// Fetch article content
|
||||||
|
console.log('📥 Fetching article content...');
|
||||||
|
const rawHTML = await fetchArticleContent(articleUrl);
|
||||||
|
const articleContent = cleanHTML(rawHTML);
|
||||||
|
console.log(`✅ Fetched ${articleContent.length} characters of content`);
|
||||||
|
|
||||||
|
// Extract data using AI
|
||||||
|
console.log('🧠 Analyzing article with Claude...');
|
||||||
|
const extractedData = await callAnthropicAPI(articleContent, articleUrl);
|
||||||
|
console.log(`✅ Extracted data for: "${extractedData.title}"`);
|
||||||
|
|
||||||
|
// Validate extraction
|
||||||
|
console.log('\n📊 Extracted data summary:');
|
||||||
|
console.log(` Title: ${extractedData.title}`);
|
||||||
|
console.log(` Slug: ${extractedData.slug}`);
|
||||||
|
console.log(` Topics: ${extractedData.topics.join(', ')}`);
|
||||||
|
console.log(` Claims: ${extractedData.claims.length}`);
|
||||||
|
console.log(` Claim Titles: ${extractedData.claimTitles.length}`);
|
||||||
|
console.log(` Key Points: ${extractedData.keyPoints.length}`);
|
||||||
|
console.log(` Statistics: ${extractedData.keyStatistics.length}`);
|
||||||
|
|
||||||
|
// Add to claims.ts
|
||||||
|
console.log('\n💾 Adding to claims.ts...');
|
||||||
|
addClaimToDataFile(extractedData);
|
||||||
|
|
||||||
|
console.log('\n✨ Successfully added new claim page!');
|
||||||
|
console.log(` View at: /claims-library/${extractedData.slug}`);
|
||||||
|
}
|
||||||
|
|
||||||
|
// CLI usage
|
||||||
|
if (require.main === module) {
|
||||||
|
const articleUrl = process.argv[2];
|
||||||
|
|
||||||
|
if (!articleUrl) {
|
||||||
|
console.error('Usage: tsx scripts/extract-article-data.ts <article-url>');
|
||||||
|
process.exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
extractAndAddArticle(articleUrl)
|
||||||
|
.then(() => {
|
||||||
|
console.log('\n🎉 Done!');
|
||||||
|
process.exit(0);
|
||||||
|
})
|
||||||
|
.catch(error => {
|
||||||
|
console.error('\n❌ Error:', error.message);
|
||||||
|
process.exit(1);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
export { extractAndAddArticle };
|
||||||
Loading…
Reference in New Issue