演講者:王奕翔教授
國立臺灣大學電機工程學系
日 期:2018年9月26日(星期三) 14:30
地 點:國立高雄大學理學院408室
講 題:On the Price of Anonymity in Heterogeneous Statistical Inference
摘 要:
Statistical inference is a fundamental task in data science, where a decision maker aims to determine a hidden parameter based on the data it collects, as well as how the probability distribution of the data depends on the target parameter. In many modern applications such as crowdsourcing and sensor networks, data is heterogeneous and collected from various sources, each of which follows different distributions. These sources, however, may be anonymous to the decision maker due to considerations in identification costs and privacy. Since the distribution becomes unknown, it is unclear what is the impact of such anonymity on the performance of statistical inference, and how to carry out optimal inference. In this talk, I will present our recent work towards settling this question, focused on binary hypothesis testing. Considering the anonymity of data sources, it is natural to formulate it as a composite hypothesis testing problem. First, we propose an optimal test called mixture likelihood ratio test, a andomized threshold test based on the ratio of the uniform mixture of all the possible distributions under one hypothesis to that under the other hypothesis. Second, we focus on the Neyman-Pearson setting and characterize the error exponent of the worst-case type-II error probability as the dimension of data tends to infinity while the proportion among the dimensions of different data sources remains constant. It turns out that the optimal exponent is a generalized divergence between the two families of distributions under the two hypotheses. Our results elucidate the price of anonymity in heterogeneous hypothesis testing and can be extended to more general inference tasks.
This talk is based on the joint work with my student Wei-Ning Chen.