edisclosure

Mapping Data Sources for Disclosure and Litigation Preservation

21st January 2026  |  6 min read

Author: Will Lunt, Commercial Director at CYFOR Legal

When legal teams think about preservation, the focus often starts with issuing a legal hold. But one of the most important early steps happens before collection and before review: mapping the data sources that may hold relevant information.

If you do not know where the data lives, it becomes much harder to preserve it properly. Relevant material can sit far beyond email, spread across shared drives, chat platforms, mobile devices, cloud applications, databases, archived systems, and third-party tools. A defensible preservation process depends on identifying those locations early and understanding how the data is stored, who controls it, and whether it is at risk of deletion.

What does “mapping data sources” mean?

Mapping data sources means creating a clear picture of:

  • where potentially relevant data is stored,
  • who uses and controls those systems,
  • how long the data is retained,
  • whether it can be changed, deleted, or overwritten,
  • and what needs to happen to preserve it.

This is not just a technical exercise. It is a practical legal and operational step that helps teams make informed decisions about scope, risk, and preservation priorities.

Why it matters

A data source map helps legal teams:

  • identify relevant data beyond the obvious locations,
  • reduce the risk of losing material through routine deletion,
  • focus preservation efforts on the right systems and custodians,
  • work more effectively with IT, compliance, and records teams,
  • and build a defensible record of the preservation steps taken.

In short, it helps move preservation from assumption to evidence.

Start with the issues, not the systems

Before listing platforms, start with the matter itself. Ask:

  • What are the issues in dispute?
  • What time period is relevant?
  • Which people, teams, or business functions are likely to hold relevant data?
  • Which transactions, projects, or decisions are central to the matter?

This gives context to the mapping exercise. Without it, teams risk creating a long inventory of systems without understanding which ones actually matter.

The key categories to map

A useful data source map should usually include the following categories:

  1. User communications
    Email, calendars, Teams chats, Slack, SMS, WhatsApp, and other direct communications.
  2. User-held documents
    Local folders, laptops, desktops, home drives, OneDrive, Google Drive, and shared folders.
  3. Team and enterprise repositories
    Document management systems, SharePoint sites, shared drives, knowledge repositories, case management tools, and collaboration spaces.
  4. Structured systems
    CRMs, ERPs, HR systems, finance platforms, ticketing systems, and other databases that may hold relevant records.
  5. Mobile and device data
    Company phones, tablets, BYOD arrangements where relevant, call logs, voicemails, and app-based communications.
  6. Cloud and SaaS platforms
    Project tools, e-signature tools, customer systems, collaboration apps, and specialist line-of-business platforms.
  7. Archived, legacy, and backup sources
    Email archives, retired systems, backups, legacy databases, and historical storage environments.
  8. Third-party held data
    Outsourced providers, consultants, payroll vendors, hosting partners, or other external organisations holding data under the client’s control or access rights.

What to record for each source

For each data source, record enough information to support preservation decisions. That usually includes:

  • System or source name
  • Description of the data held
  • The likely relevance to the matter
  • Named custodian or owner
  • System owner or IT contact
  • Date range available
  • Retention period
  • Auto-delete or overwrite risks
  • Export/preservation options
  • Access restrictions and any known limitations

This turns a simple list into something operationally useful.

Questions legal teams should ask

When mapping data sources, some of the most useful questions are:

  • Who created, received, or managed the relevant information?
  • Where would they normally store or discuss it?
  • Are there shared repositories beyond personal accounts?
  • Are chat messages or mobile communications likely to be relevant?
  • Does any data sit in systems with short retention periods?
  • Are any relevant accounts linked to former employees?
  • Is any data held by third parties or in specialist business systems?
  • Are there ongoing migrations, refreshes, or decommissioning projects that could affect the data?

These questions often reveal risk areas that a standard hold notice alone will not catch.

Commonly missed data sources

Some of the most frequently overlooked areas include:

  • Teams or Slack chats,
  • personal and shared cloud storage,
  • former employee mailboxes and accounts,
  • mobile devices and app-based messaging,
  • project management tools,
  • third-party SaaS platforms,
  • archived data,
  • and systems with automatic deletion or short retention settings.

These sources are easy to miss because they sit outside the traditional “email and folder” mindset.

A practical approach to mapping data sources

A sensible approach is usually:

Step 1: Define the scope
Clarify the issues, date range, business units, and likely custodians.

Step 2: Speak to the right people
Bring together legal, IT, records, compliance, and relevant business stakeholders.

Step 3: Identify likely repositories
List the systems, devices, and platforms where relevant information may sit.

Step 4: Assess preservation risk
Flag sources with auto-delete, overwriting, access limitations, or ongoing operational change.

Step 5: Prioritise action
Not every source carries the same risk. Focus first on the places where relevant data is both likely and vulnerable.

Step 6: Document decisions
Record what was included, what was excluded, and why.

The goal is defensibility, not excess

Mapping data sources is not about preserving every system in the business. It is about taking reasonable, informed steps to identify where relevant material may exist and acting before it is lost.

A clear data source map helps legal teams preserve more effectively, communicate better with internal stakeholders, and reduce avoidable risk later in the disclosure process.

Secret Link