Python Automation & Web Scraping: Production Pipelines & Cloud Daemons

Level: Beginner to Intermediate 16 total hours DBERT Platform

Course curriculum

Day 1: Day 1: Modern Python Scripting, File I/O & OS-Level Automation 3 subtopics
▸
Pathlib, Cross-Platform File Systems & Directory Walks

Upgrade your filesystem automation skills. Replace legacy, fragile os.path string concatenations with object-oriented pathlib.Path. Navigate directory structures safely across Windows and Linux. Perform recursive file discovery with rglob(), batch file renaming, and extract file metadata (size, created/modified timestamps). Build automated file tidying scripts that organize chaotic directories by file extension and date stamps.

▸
Structured Data Parsing: JSON, CSV, YAML & Config Management

Parse and manipulate diverse structured data feeds. Ingest and serialize nested JSON objects with proper encoding handling. Stream multi-gigabyte CSV transaction exports row-by-row using csv.DictReader to prevent server memory exhaustion. Parse YAML configuration files for automation settings, and isolate sensitive credentials using python-dotenv with strictly scoped environment variables.

▸
Subprocess Management & Shell Command Orchestration

Command operating system utilities directly from Python. Use subprocess.run and subprocess.Popen safely: capture stdout and stderr streams, enforce non-zero exit code validation (check=True), and set execution timeouts to prevent hung processes from blocking servers. Eliminate critical shell injection vulnerabilities by passing command arguments as safe argument lists rather than raw strings with shell=True.

Day 2: Day 2: Web Scraping & HTML Parsing with BeautifulSoup & Requests 3 subtopics
▸
HTTP Client Engineering: Sessions, Headers & Resilient Retries

Build industrial-strength HTTP scrapers. Use requests.Session to maintain persistent TCP connections (HTTP Keep-Alive) and manage cookies across requests. Configure realistic browser headers (User-Agent, Accept, Accept-Language) to avoid naive bot blockers. Mount custom HTTP adapters with urllib3.util.retry.Retry to implement automatic exponential backoff retries when facing 429 Too Many Requests or intermittent 502/503 server errors.

▸
DOM Parsing, CSS Selectors & XPath with BeautifulSoup

Extract structured data from raw HTML trees. Compare parsing engines (lxml vs html.parser). Master concise CSS selector syntax with soup.select() and soup.select_one(). Navigate complex DOM hierarchies using relative traversal: parent, next_sibling, and find_previous. Extract link hrefs, image sources, and parse nested tabular data into clean Python dictionaries and lists.

▸
Anti-Bot Defense Navigation: Headers, Proxies & Rate Governance

Scrape sustainably and responsibly. Navigate modern rate limiters and Web Application Firewalls (WAF). Implement randomized jitter delays (random.uniform) between requests to avoid predictable bot cadence. Configure HTTP and SOCKS5 proxy pools to distribute outbound traffic. Parse and honor target site robots.txt files, identify honeypot links, and manage ethical scraping boundaries.

Day 3: Day 3: Browser Automation & Dynamic SPAs with Playwright & Selenium 3 subtopics
▸
Headless Browser Architecture & Playwright Setup

Automate modern dynamic web applications (React, Angular, Vue) where target data is rendered asynchronously via client-side JavaScript. Understand headless browser architecture using the Chrome DevTools Protocol (CDP). Set up Playwright in Python on headless Linux EC2 servers without physical displays. Configure browser contexts, viewport sizes, and optimize memory footprints by blocking unnecessary image and font downloads.

▸
Dynamic Form Interaction, Infinite Scroll & Explicit Waits

Interact with dynamic UI elements reliably. Learn why time.sleep() causes fragile, slow automation. Master Explicit Waits (wait_for_selector, wait_for_load_state) that wait for precise DOM conditions. Automate multi-step form submissions, dropdowns, modal dismissals, and infinite scroll pagination by evaluating JavaScript scroll positions until content loading terminates.

▸
Automated Screenshots, PDF Generation & File Interception

Capture visual and binary digital assets. Generate full-page screenshots of dashboards for visual monitoring and archival. Print dynamic web views into high-resolution, print-ready PDF reports. Automate binary file downloads (CSV reports, invoices), intercept download events with Playwright, and verify downloaded file integrity using MD5 checksums.

Day 4: Day 4: Data Pipeline Automation with Pandas & Tabular Processing 3 subtopics
▸
Vectorized Data Transformation & ETL Pipelines

Process structured automation feeds at scale. Ingest messy tabular datasets into Pandas DataFrames. Implement vectorized data cleaning: handling missing null values (fillna, dropna), removing duplicate records, regex string standardization (formatting phone numbers, stripping currencies), type casting, and parsing heterogeneous datetime formats into standardized UTC timestamps.

▸
Excel Automation & Report Formatting with OpenPyXL

Automate corporate spreadsheet workflows. Read, create, and modify .xlsx workbooks using openpyxl. Populate multi-sheet workbooks, inject native Excel formulas (=SUM, =AVERAGE, =XLOOKUP), apply custom cell formatting (currency formats, font styling, borders, fill colors), configure conditional formatting rules, and auto-adjust column widths to fit content perfectly.

▸
Automated PDF Extraction & Document Processing

Extract structured data from unstructured corporate documents. Use pdfplumber and pypdf to inspect digital PDF invoices, statements, and receipts. Extract clean raw text, extract bounding box coordinates, and reconstruct complex tabular data grids into clean Pandas DataFrames. Automate text normalization across multi-page document archives.

Day 5: Day 5: API Integration, Webhooks & Automated Messaging Systems 3 subtopics
▸
Consuming Third-Party REST APIs & OAuth2 Workflows

Interface with external cloud services. Manage API authentication: Static API keys, Bearer tokens, and the OAuth2 Client Credentials flow. Handle cursor-based and offset-based pagination across long API collections. Inspect and respect rate-limiting headers (X-RateLimit-Remaining, X-RateLimit-Reset) and cache idempotent GET responses locally to avoid API quota burn.

▸
Multi-Channel Alert Pipelines: Email (SMTP/cPanel), Slack & Telegram

Notify humans when critical events occur. Build multi-channel alerting engines: Dispatch rich HTML emails with embedded styles and PDF/Excel attachments using Python's smtplib and email.mime modules or transactional email APIs (cPanel/SES). Broadcast structured JSON alert webhooks to Slack channels (Block Kit formatting) and Telegram messaging bots.

▸
Building Lightweight Webhook Receivers with Flask

Create inbound event listeners that trigger automation pipelines. Build a lightweight Flask webhook endpoint to receive inbound event notifications (e.g. Stripe payment succeeded, GitHub push, Typeform submission). Validate HMAC cryptographic signatures (SHA-256) to verify webhook authenticity and reject spoofed payloads. Return immediate 200 OK responses before dispatching asynchronous background tasks.

Day 6: Day 6: Scheduled Execution, Background Jobs & Process Daemons 3 subtopics
▸
Linux Cron Jobs & Systemd Timers

Automate recurring scheduled tasks on Linux servers. Understand the 5-field cron syntax (minute hour day month day-of-week). Troubleshoot classic cron failures: missing PATH environment variables, wrong working directories, and lost error logs. Transition from legacy cron to modern Linux Systemd Timers for centralized logging in journalctl, dependency chaining, and persistent miss handling across reboots.

▸
In-Process Schedulers: APScheduler & Background Workers

Embed scheduling directly inside Python applications. Use APScheduler (Advanced Python Scheduler) for interval, date, and cron triggers. Configure persistent job stores (SQLite or PostgreSQL) so scheduled jobs survive process restarts. Configure misfire grace times, coalescing missed executions, and understand concurrency trade-offs between thread pools and process pools.

▸
Concurrency Control, File Locks & Idempotency

Prevent catastrophic race conditions in automated workflows. What happens when a 5-minute scheduled scraper takes 12 minutes to run? Prevent overlapping concurrent executions using file locks (fcntl.flock on Linux, portalocker) and PID tracking files. Design idempotent database mutations that safely handle repeat executions without creating duplicate records or double-charging accounts.

Day 7: Day 7: Production Hardening, Logging, Error Alerting & EC2 Deployment 3 subtopics
▸
Enterprise Logging Architecture & Rotating Handlers

Eliminate unmaintainable print() debugging statements. Configure Python's standard logging module for production: Log levels (DEBUG, INFO, WARNING, ERROR, CRITICAL), timestamp formatting, module attribution, and process IDs. Configure rotating file handlers (RotatingFileHandler by size, TimedRotatingFileHandler by day) to prevent log files from exhausting EC2 disk storage. Structure logs in JSON format for automated ingestion.

▸
Self-Healing Automation: Retries, Dead Letter Queues & Health Checks

Design automation pipelines that survive chaos. Build custom retry decorators with jittered exponential backoff. Implement failure isolation: an exception processing record #42 should not crash the batch of 1,000 records. Route failed records into an Error Dead Letter Queue (quarantine table or file) with full traceback details for later inspection. Implement heartbeat monitoring pings (e.g. Healthchecks.io, BetterStack) to detect silent task death.

▸
Headless Linux EC2 Deployment & Systemd Service Management

Deploy headless automation daemons to AWS EC2. Structure application code under /var/www/apps/automation. Configure virtual environments, isolate environment secrets, and assign restricted non-root service user ownership. Write and enable persistent Linux Systemd service unit files with Restart=always, manage automatic boot activation, and monitor real-time execution logs with journalctl.

Enrollment

Access type Free quota
Capstone project Required (DBERT verified)