About this role
Phone numbers and emails in this ad are masked until you log in.
auto_translated_note
Programmatic QA • Testing for LLMs & Agents • Data Quality • Platform Reliability About Finalytics.ai Finalytics.ai is the leading provider of personalization for the financial industry. Our platform combines data integrations, machine learning, and real-time technology to make digital experiences more relevant and higher-converting for credit unions and banks. We're a growing startup led by industry veterans, building the next generation of AI-driven personalization.
Why This Role Is Different QA at Finalytics goes well beyond clicking through a UI. Our platform makes model-driven decisions, runs LLMs and agents that generate content and answer questions, and depends on data pipelines that feed those models every day - and all of it has to be tested programmatically. We're looking for an engineering-minded QA team contributor to help build quality across three areas: our core personalization features, our LLM and agentic capabilities, and the data that powers them.
This is a coding role, embedded in the same repo and release flow as our engineers that will report directly to the CTO. You won't just find bugs - you'll build the automated tests, evals, and data checks that let a small team ship trustworthy AI every sprint. Our stack is Python/Django with a JavaScript personalization tag, backed by MySQL, Celery, BigQuery, and AWS.
What You'll Do 1. Programmatic QA of Core Features - Extend our scenario test runner - a proprietary harness that captures real production personalization requests and replays them across environments, asserting on expected algorithms and content selection. Grow it into automated regression across every client. - Write automated tests in Python with pytest across our tiers - unit, integration, HTTP, and end-to-end. - Build headless Playwright end-to-end tests to verify how personalized content and tracking render on real client pages. - Harden the pre-deploy quality gate and pre-commit checks that block bad changes automatically.
2. Testing & Standardizing LLMs and Agents - Design evals for non-deterministic AI features - our conversational analytics assistant, AI content builders, and generative SEO - measuring correctness, grounding, and regression across prompt and model versions. - Test the tool-calling and agentic layers - that function-calling loops pick the right tools and guardrails hold on adversarial input. - Validate our agent/MCP interface - contract conformance, rate limiting, authorization, and safe failure. - Help set our standards for shipping AI - catching hallucinations and drift, and benchmarking prompt/model changes before clients see them. 3.
Data Quality Engineering - Build automated data-health checks that flag stale rollups, incomplete coverage, and broken aggregations before they hit a client dashboard. - Validate data pipelines end-to-end - rollups, funnel/rate/financial ingestion, and BigQuery - with drift detection across environments. - Guard model inputs so the signals our ML depends on stay accurate and complete. 4. Reliability & Performance - Track platform performance - response times, JS load, and page speed - and help keep it fast. - Stand up quality dashboards - uptime, coverage, data-health, and eval scores.
5. Collaboration & Bug Lifecycle - Work in the codebase alongside engineers to diagnose issues across development, release, and deployment. - Drive the bug lifecycle - reproduce, capture with a failing test, and verify the fix.
What We're Looking For
- 3+ years in QA/SDET or test automation with a code-first approach. - Strong Python - you write clean test code and can read the app you're testing. - pytest (preferred) and browser automation (Playwright or Selenium). - API and contract testing experience. - A genuine interest in testing AI - comfortable with non-determinism, evals, and prompts. - Data-savvy - strong SQL, and the instinct to validate pipelines and reconcile data. - Building automated quality gates into the deploy and release process. Nice to Have - Testing or evaluating LLM applications - evals, prompt regression, tool-calling agents, or MCP. - Data or analytics QA - BigQuery or ETL/rollup validation. - Django, MySQL, or Celery experience. - Security testing with SAST/DAST tooling. - Familiarity with machine learning. - Financial industry, personalization, or CMS/marketing-platform experience. - Familiarity with AWS. - SaaS startup experience on a fast-moving, multi-tenant platform. Why Finalytics - Frontier work - help define what QA means for AI, agents, and data-driven personalization in finance. - Direct impact - help shape how quality works across the platform, reporting straight to the CTO. - Automation-first culture - your work is code, in the same repo and release flow as engineering. - Remote-first, collaborative, low-ego team growing with a scaling fintech. Apply directly on RemoteJobs.org: https://remotejobs.org/remote-jobs/qa-engineer-sdet-ai-data-platform-quality-extractable
Community Q&A
Anyone worked here? Ask before you apply.
No threads yet for this job or company.