Weekly Fivetran audit of sensitive data in your warehouse

By General Input

Every Monday we review what your Fivetran pipelines are copying into the warehouse, flag unmasked personal data in Slack, and open a ticket for each affected pipeline.

Integrations

  • Fivetran
  • Slack Bot
  • Linear

Type

Agentic Task

Categories

  • Operations
  • Engineering

Every Monday at 8am, review what my Fivetran pipelines are actually replicating into the warehouse and flag sensitive data that is landing in the clear. This is a governance sweep over pipeline configuration, not a sync health check, so I care about what is being copied rather than whether the last run succeeded.

Start with List Connections to get every connection in the account, then walk them one at a time. For each connection, use Get Connection Schema Config to inspect the full schema, table and column configuration, and determine which tables and columns are enabled for syncing and which are hashed or blocked. Where the schema config does not give enough column level detail for a table, follow up with Get Source Table Columns Config for that specific source table.

Judge each enabled column by its name together with its table and schema context to identify likely regulated or sensitive fields. Treat the following as sensitive: email addresses, phone numbers, date of birth, national ID or SSN, passport and driver licence numbers, bank account and payment card numbers, salary and compensation, and free text health or notes fields. Context matters, so a column called number inside a payment_methods table is a very different risk from number inside a line_items table. Use the table and schema name to disambiguate before you decide.

The finding I care about is a column that is enabled AND not hashed. If a sensitive column is already hashed, or already blocked or disabled, it is fine and should be left out of the report entirely. Suppressing already remediated columns is what keeps the Monday message readable. Do not attempt to diff against previous runs or carry state between weeks, just report the current state of the configuration each time.

Post one Slack Bot message with Send a Message to my data governance channel. Group findings by connection, and order them so the highest risk leads: national ID and SSN, payment card and bank details, and free text health or notes first, then date of birth and passport, then email, phone and salary. For every finding, name the exact schema, table and column so someone can go straight to it without hunting. Open the message with a one line summary of how many connections were reviewed, how many had findings, and how many were clean.

For each connection with at least one finding, use Create Issue in Linear to open one issue on my data team. Title it with the connection name, list every affected column in schema.table.column form, and give each one a recommended fix of either hash or disable. Recommend hash where the field still has analytical value in masked form, such as email or phone used for joins and counts, and recommend disable where the warehouse has no business holding the value at all, such as a passport number or a full card number. Keep it to one issue per connection so each ticket maps to a single owner and a single pipeline.

Do not change any pipeline configuration by default. This workflow is advisory. The single exception is Modify Column Config, which you may use to hash a column only when that column name matches a pattern I have explicitly pre-approved in advance. My pre-approved patterns are: FILL THIS IN, for example columns named exactly email or phone_number. If that list is empty, change nothing at all. Never disable a column automatically, never hash anything outside the approved patterns even when the risk looks obvious, and never modify table or schema level config.

Always close the Slack message by stating plainly what was changed versus what still needs a human decision. If anything was hashed automatically, list it under a Changed automatically heading along with the approved pattern that authorized it. Put everything else under a Needs your decision heading with a link or reference to the Linear issue that tracks it. If nothing was changed, say that explicitly rather than staying silent, so I never have to guess whether the workflow touched production config.

Related prompts

Explore more prompts
A brand asset library your marketing team actually searchesTurn Mailjet email clicks into ranked HubSpot follow-upsClean out the Looker dashboards and Looks nobody opensLiveKit live operations console for room moderationWake up dormant Keap leads with a researched reasonLiveChat coverage board for planning next week's shiftsPhone routing control panel for LiveKit voice agentsLinkedIn Ads budget pacing dashboard for every client accountGive your team Looker numbers without buying more seatsPause marketing emails to escalated customers, then restore them