This work model a preparation plan as a sequence of deterministic curation operators and asks how each step changes the evidence available to an observer with prior knowledge about a target, which is an inference-aware foundation for guiding curators throughout data preparation.
Abstract
Data preparation often begins with sensitive data and produces a releasable artifact for analysis, sharing, or model training. Existing workflows are primarily guided by utility: a curator drops attributes, coarsens values, filters populations, and suppresses tuples until the resulting dataset appears useful and safe. Privacy, when considered, is usually evaluated only on the final release. We propose privacy-aware data preparation as an interactive guidance problem. We model a preparation plan as a sequence of deterministic curation operators and ask how each step changes the evidence available to an observer with prior knowledge about a target. Our semantics is based on compatibility sets, which capture the source tuples still plausible for the target after a released representation is observed. This view separates operators that remove evidence from those that remove ambiguity, explains why privacy effects can be non-monotone, and supports prefix-level feedback under a disclosure budget. The result is an inference-aware foundation for guiding curators throughout data preparation, rather than judging privacy only after the final artifact is produced. We conclude by identifying the key challenges in building interactive, inference-aware data preparation systems.
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.
ProxyDrift is presented, a framework that identifies and measures drift between production traffic and offline evaluation sets, and constructs and refreshes those evaluation sets accordingly; all without access to raw user data.
Michael Levit, Josh Ledgard, Haoyu Dong et al.· 0 citations
This work investigates a probabilistic variant of PCD, where an LLM-driven probabilistic estimation of k-anonymity is augmented with an LLM-driven probabilistic estimation of k-anonymity, and proposes k-anonymity as a useful auxiliary metric for tackling PCD.
Sharing databases and data streams imposes the danger of revealing private information in the form of complex events which can comprise individual data elements and their combinations. Identifying these privacy-revealing complex events is crucial for preserving privacy while maintaining data utility. However, data producers often lack the expertise to comprehensively identify these events, which undermines many state-of-the-art privacy-preserving mechanisms that rely on accurate event labeling. To address this challenge, we developed pArborist - a tool that can semi-automatically create a set of queries to identify and label privacy-revealing complex events in both static datasets and dynamic data streams, guided by the privacy requirements of the data producer. pArborist uses the schema of the database or data stream combined with initial input from the data producer, i.e., seed queries. From each seed query, pArborist grows a tree containing all possible syntactically correct queries, constrained by an upper limit on computational resources. Following this growing phase, the tree is refined by eliminating queries that lack correlation to the seed or are conditionally independent of the seed. Our evaluation indicates that pArborist achieves overall recall of \(90\%\) and precision of \(93\%\) in finding privacy-revealing queries, and this significantly surpasses the state-of-the-art approach FQID. In data stream processing experiments, pArborist introduces a delay of approximately 1.3 ms following an average warm-up period of 920 ms. The experiments also show that pArborist can automatically detect privacy-revealing complex events according to GDPR.
He Gu, Thomas Plagemann, V. Goebel· International Conference on...· 0 citations
The publication of government data promotes transparency and supports advances in scientific research. Such publications may occur either mandatorily or upon request. However, published datasets may contain sensitive information about data subjects, potentially leading to privacy concerns. To address these issues, many countries have established regulations governing the handling of personal data. Nevertheless, these regulations generally do not specify the technical mechanisms that data managers should adopt when publishing data. The literature proposes several privacy-preserving models that can be employed to protect sensitive information during data publication. However, the application of these models often reduces data utility and may even render the published data unsuitable for certain purposes. Consequently, achieving an appropriate balance between privacy preservation and data utility remains a major challenge in data publication. This article proposes a methodology for privacy-preserving publication of structured tabular data under two scenarios: mandatory publication, in which the data manager is legally required to disclose the data, and requested publication, in which data are released upon request. In the requested publication scenario, the methodology introduces an interactive process through which data managers and requesters collaboratively negotiate the transformations applied to the data in order to balance privacy and utility. To evaluate the proposed methodology, a case study was conducted using a real-world dataset, including an assessment based on the Technology Acceptance Model (TAM). The results indicate that the methodology supports privacy-preserving data publication and enables a collaborative balance between privacy and utility through the proposed interactive process.
Bruno Roberto Silva De Moraes, Ariel Soares Teles, Josenildo Costa da Silva et al.· IEEE Access· 0 citations
With the rapid growth of mobile applications, user data privacy has become an increasing concern. While privacy policies describe how apps collect and share data, platforms such as Google Play provide Data Safety labels intended to summarize these practices. Because these disclosure channels are declared separately, they may present inconsistent representations of app data practices, creating uncertainty for users and regulators. In this work, we conducted a large-scale empirical study of disclosure consistency across 6,051 Android apps. Using an LLM-based extraction framework and a unified schema over 14 Google Play data categories and two operations (collection and sharing), we measure per-app and per-category consistency and introduce a sensitivity-weighted risk score that emphasizes high-risk data types. We find that misalignment disproportionately affects sensitive categories such as personal information and device identifiers, with sharing disclosures exhibiting lower consistency than collection disclosures. Elevated privacy risk is concentrated in app categories associated with persistent monitoring and communication. Overall, our findings highlight structural gaps in current disclosure mechanisms and underscore the need for stronger verification and greater transparency in platform-level privacy reporting.
Mst. Eshita Khatun, Lamine Noureddine, Sideeq Bello et al.· Proceedings on Privacy Enhan...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.