Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.
Role SummaryOwn and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities- Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
- Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
- Introduce proxy rotation and egress management; retire the single-IP failure mode.
- Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
- Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
- Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
- Containerize and schedule the pipeline; add CI running the offline tests on every change.
- Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]
Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]
Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]
Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]
Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]
Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]
Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]
Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]
Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]
Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]
Experience — Required- 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
- Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
- Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
- Mentored engineers; set crawl standards, review practice, and on-call runbooks.
- Degree optional — equivalent practical experience is fully accepted.
- US payer policy, formulary, or prior-authorization document domain knowledge.
- Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
- LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/.
- Compliance or legal-review exposure on data acquisition programmes.
Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Skills Required
- 8+ years of experience in data acquisition
- 5+ years owning a web scraping platform end to end
- Production experience covering 1,000+ distinct domains or 10M+ pages per month
- Experience rescuing a brittle legacy scraper and demonstrating reliability and cost improvements
- Experience mentoring engineers and establishing crawl standards, review practices, and on-call runbooks
- Expert Python skills, including asynchronous programming, typing, and profiling
- Advanced web protocol, browser automation, crawling, reverse engineering, extraction, anti-bot, data engineering, testing, and observability experience
- Knowledge of legal and ethical data acquisition, including robots.txt, rate limits, terms of service, CFAA, GDPR, PII minimization, and licensed APIs
- US payer policy, formulary, or prior-authorization document knowledge
- Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library
- Experience with cost-controlled LLM-assisted extraction and AWS Bedrock
- Compliance or legal-review experience related to data acquisition programs
- A degree is optional; equivalent practical experience is accepted
What We Do
Proof-of-Skill is a hiring and talent-matching platform that helps candidates find roles aligned with their skills, values, experience, and goals. It verifies identity, education, work history, and skills through proctored assignments and expert validation, then connects verified candidates with relevant employers and interviews. For businesses, it provides candidate skill-assessment software intended to reduce resume screening and improve recruiting efficiency for modern teams and hiring decisions.






