- Update base_scraper.py convert_to_markdown() to properly clean HTML - Remove script/style blocks and their content before conversion - Strip inline JavaScript event handlers - Clean up br tags and excessive blank lines - Fix malformed comparison operators that look like tags - Add comprehensive HTML cleaning during content extraction (not after) - Test confirms WordPress content now generates clean markdown without HTML This ensures all future WordPress scraping produces specification-compliant markdown without any HTML/XML contamination. |
||
|---|---|---|
| .. | ||
| .cookies | ||
| .sessions | ||
| backlog | ||
| debug/.sessions | ||
| recent | ||
| wordpress_clean | ||
| test_wordpress.md | ||
| test_youtube.md | ||
| tiktok_advanced_test.md | ||
| wordpress_content.html | ||
| wordpress_content.md | ||
| wordpress_markdownify.md | ||
| wordpress_post_raw.json | ||