Data handling
This page is for procurement, security teams, and anyone who needs to answer “what does Marriska actually do with our data?” before signing off on the platform.
For the auth/access side specifically (sessions, API keys, role checks), see Security model.
What we store
Section titled “What we store”About the test itself
Section titled “About the test itself”- Test definitions: the natural-language description you wrote, the parsed steps, any variables, the target URL.
- Test sets: groupings of tests, names, ordering.
- Schedules: cron expression, timezone, browser selection, notification recipients.
- Tags, projects, folders: organizational metadata.
About each run
Section titled “About each run”- Step results: pass/fail, duration in milliseconds, error message (when failed).
- Step screenshots: stored as
LargeBinaryblobs in Postgres (columnscreenshot_dataontest_run_results). One per step per browser. - Visual baselines: when you opt into visual regression, the baseline screenshot is stored against the step ID + browser.
- Logs: per-step structured log entries (
step_log) with the action, target, URL at the time of the step.
Values from variables you mark sensitive are masked (••••••) in
step results, descriptions, logs, and the per-iteration values before
they’re stored — the real value is still typed into the page, only what
we record is hidden. Screenshots box out form fields holding a secret at
capture (only a secret rendered as page text can still appear). See
Mask passwords and other secrets.
About people and orgs
Section titled “About people and orgs”- Users: email, display name, password hash (bcrypt) for email/password accounts; provider-issued IDs for OAuth (Google, GitHub).
- Organizations: name, plan tier, max member count.
- Memberships: who’s in which org and at what role (owner / admin / member / viewer).
- API keys: stored as SHA-256 hashes, never the raw key. The
prefix (
ak_live_••••••••...abcd) is kept for display. - Sessions: tokens are hashed; we store the device label and the last-used time.
Where data lives
Section titled “Where data lives”- Database: PostgreSQL 16. Test definitions, runs, screenshots, baselines, users, sessions — all in one Postgres cluster.
- Screenshots and baselines specifically: in Postgres
LargeBinarycolumns, not on a separate object store. (For the curious: this is intentional — backup, retention, and access control all collapse to one system.) - In-flight execution state: in process memory on the executor, cleaned up when the run finishes.
- Self-hosted option: every byte of the above can run on your
infrastructure if you prefer (
STORAGE_MODE=databaseagainst your own Postgres). Enterprise tier includes the self-hosted option.
There’s an InMemory storage mode (STORAGE_MODE=memory) used for
development and ephemeral testing — runs vanish on backend restart.
Production uses STORAGE_MODE=database.
What gets sent to AI providers
Section titled “What gets sent to AI providers”Each AI task ships only what it needs to do its job:
| Task | Sent to provider |
|---|---|
| Translation (non-English description → English) | The natural-language text you wrote (can be routed to a local Ollama model — no external send — depending on configuration) |
| Parsing (English → structured steps) | The English text + a static system prompt |
| Visual comparison | Two PNG screenshots (baseline + current), encoded as base64. Tests with sensitive variables are excluded — their screenshots never go to the vision model (local pixel-diff only). For other tests, screenshots aren’t redacted, so a secret visible on the page can be in the image — see sensitive variables |
| Generation (when you ask for a generated test) | The prompt you provided |
Nothing else from the run leaves the platform on these calls. Step results, screenshots from non-visual steps, user identity, org name — none of that goes to the AI provider.
When BYOK is configured (per-user BYOK is built but disabled unless the
operator sets BYOK_ENCRYPTION_KEY — see BYOK),
these calls go to your provider account instead of ours. That puts
the inference logs in your provider’s dashboard rather than ours.
What gets logged
Section titled “What gets logged”We log:
- Auth failures and successes (without the password or token).
- API key usage — the key’s hashed prefix and last-used timestamp, not the raw value.
- Per-request metadata — endpoint, status code, duration.
- Sentry errors if
SENTRY_DSNis configured (production).
We don’t log:
- Passwords. Ever.
- Raw API keys. Tokens are masked after the first 8 characters in any log line that touches them.
- Test page contents (HTML/DOM dumps). The browser screenshot and the step result are persisted; the page source isn’t.
Retention and deletion
Section titled “Retention and deletion”History retention by tier
Section titled “History retention by tier”Run history is retained per-tier:
- Free: 14 days
- Starter: 90 days
- Pro: 180 days
- Team: 2 years
- Enterprise: unlimited
The history_retention_days value is enforced two ways: a nightly
background job prunes runs older than your tier’s window, and a
downgrade prunes on the spot (the impact preview shows what shrinks).
The nightly purge runs on the production database deployment — it
doesn’t fire in single-worker in-memory dev mode.
Deleting a test or test set
Section titled “Deleting a test or test set”Deleting a test removes the definition and its variable data. Run history for past executions is retained per the retention policy above.
Exporting your data
Section titled “Exporting your data”Download everything you’ve authored from Settings → Account → Danger zone → Export data — a JSON file with your profile, organizations, projects, folders, tests, test sets, schedules, and tags. Run history and screenshots aren’t in the export (they age out per the retention window above).
Deleting an account
Section titled “Deleting an account”Delete your account yourself from Settings → Account → Danger zone; type your email to confirm. It’s a soft delete with a 30-day grace period — log back in within 30 days to cancel it. After that, a nightly job hard-deletes your account and the data that cascades from it (projects, tests, tags, test sets, runs, sessions). You can’t delete while you’re the sole owner of a shared org that still has other members — transfer or remove them first. See Export or delete your account.
Sessions and tokens
Section titled “Sessions and tokens”- Active sessions show under Settings → Security → Active Sessions. Revoke individual sessions or sign out everywhere.
- Email verification links expire after 15 minutes (
EMAIL_VERIFICATION_EXPIRE_MINUTES). - Invite links expire after 7 days.
Who can see what
Section titled “Who can see what”- Within an org: members see what their role allows. Viewers can read; members can edit and run; admins/owner can manage settings, members, and API keys.
- Across orgs: every database read filters by the caller’s
org_id. Cross-org reads return403. The middleware enforces this on every protected route. - Reports: today, every report URL requires the recipient to be a member of the same org. Public share tokens aren’t shipped — see Sharing reports.
Encryption
Section titled “Encryption”- In transit: HTTPS by default for the production API
(
https://api.marriska.com); the security headers middleware addsX-Content-Type-Options: nosniff,X-Frame-Options: DENY,X-XSS-Protection: 1; mode=block,Referrer-Policy: strict-origin-when-cross-origin, and ano-storecache directive on every/api/response. - At rest: Postgres-level encryption depends on your hosting provider (Railway / your own infra). API keys are SHA-256 hashed before storage; passwords are bcrypt-hashed.
- Per-user secrets (per-user BYOK keys): stored Fernet-encrypted
in the database. The store is built but disabled unless the operator
sets
BYOK_ENCRYPTION_KEY— see BYOK.
What we don’t do
Section titled “What we don’t do”- Sell or share data with third parties. Period.
- Train models on your tests. AI provider calls go through their standard inference APIs — not their training pipelines. (For provider-side guarantees on this, refer to your chosen provider’s data-handling docs — OpenAI, Anthropic, etc., all publish API data-use policies.)
- Store payment-card data. When real Stripe ships, card data lives in Stripe; we hold the customer ID, not the card.
Related
Section titled “Related”- Mask passwords and other secrets — hide sensitive variable values in reports and logs
- Security model — auth, isolation, RBAC, the access-control side
- BYOK — moving AI inference to your own provider account
- Plan tier limits — history retention by tier