Hansheng Wang
Dados Biográficos
| ID | 8920099 |
|---|---|
| NOME | Hansheng Wang |
| PRENOMES | Hansheng |
| SOBRENOME | Wang |
| ASSINATURA | WANG H |
| AFILIAÇÕES | Peking University |
| ORCID | 0000-0003-2386-0209 |
| VERIFICADO | Sim |
| TOTAL DE OBRAS | 19 |
| TOTAL DE CITAÇÕES | 0 |
| TOTAL COMO AUTOR | 19 |
| TOTAL COMO EDITOR | 0 |
| PRIMEIRO ANO DE PUBLICAÇÃO | 2003 |
| ANO MAIS RECENTE DE PUBLICAÇÃO | 2025 |
| ÍNDICE H | 0 |
Academic literature recommendation in large-scale citation networks enhanced by large language models
Penalized Sparse Covariance Regression with High Dimensional Covariates
Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional predictors (i.e., similarity matrices). Specifically, we use the penalization method to identify the informative predictors and estimate their associated coefficients simultaneously. We first investigate the Lasso…
Optimal Subsampling Bootstrap for Massive Data
The bootstrap is a widely used procedure for statistical inference because of its simplicity and attractive statistical properties. However, the vanilla version of bootstrap is no longer feasible computationally for many modern massive datasets due to the need to repeatedly resample the entire data. Therefore, several improvements to the bootstrap method have been made in recent years, which assess the quality of estimators by subsampling the ful…
Learning Human Activity Patterns Using Clustered Point Processes With Active and Inactive States
Modeling event patterns is a central task in a wide range of disciplines. In applications such as studying human activity patterns, events often arrive clustered with sporadic and long periods of inactivity. Such heterogeneity in event patterns poses challenges for existing point process models. In this article, we propose a new class of clustered point processes that alternate between active and inactive states. The proposed model is flexible, h…
Network Gradient Descent Algorithm for Decentralized Federated Learning
We study a fully decentralized federated learning algorithm, which is a novel gradient descent algorithm executed on a communication-based network. For convenience, we refer to it as a network gradient descent (NGD) method. In the NGD method, only statistics (e.g., parameter estimates) need to be communicated, minimizing the risk of privacy. Meanwhile, different clients communicate with each other directly according to a carefully designed networ…
Interactive Geological Data Visualization in an Immersive Environment
Underground flow paths (UFP) often play an important role in the illustration of geological data by geologists, especially in illustrating geological data and revealing stratigraphic structures, which can help domain experts in their exploration of petroleum information. In this paper, we present a new immersive visualization tool to help domain experts better illustrate stratigraphic data. We use a visualization method based on bit-array-based 3…
Feature Screening for Massive Data Analysis by Subsampling
Modern statistical analysis often encounters massive datasets with ultrahigh-dimensional features. In this work, we develop a subsampling approach for feature screening with massive datasets. The approach is implemented by repeated subsampling of massive data and can be used for analyzing tasks with memory constraints. To conduct the procedure, we first calculate an R-squared screening measure (and related sample moments) based on subsamples. Sec…
A Note on Distributed Quantile Regression by Pilot Sampling and One-Step Updating
Quantile regression is a method of fundamental importance. How to efficiently conduct quantile regression for a large dataset on a distributed system is of great importance. We show that the popularly used one-shot estimation is statistically inefficient if data are not randomly distributed across different workers. To fix the problem, a novel one-step estimation method is developed with the following nice properties. First, the algorithm is comm…
Autoregressive Model With Spatial Dependence and Missing Data
We study herein an autoregressive model with spatially correlated error terms and missing data. A logistic regression model with completely observed covariates is used to model the missingness mechanism. An autoregressive model is used to accommodate time series dependence, and a spatial error model is used to capture spatial dependence. To estimate the model, a weighted least squares estimator is developed for the temporal component, and a weigh…
Sequential Text-Term Selection in Vector Space Models
Text mining has recently attracted a great deal of attention with the accumulation of text documents in all fields. In this article, we focus on the use of textual information to explain continuous variables in the framework of linear regressions. To handle the unstructured texts, one common practice is to structuralize the text documents via vector space models. However, using words or phrases as the basic analysis terms in vector space models i…
Covariance Matrix Estimation via Network Structure
In this article, we employ a regression formulation to estimate the high-dimensional covariance matrix for a given network structure. Using prior information contained in the network relationships, we model the covariance as a polynomial function of the symmetric adjacency matrix. Accordingly, the problem of estimating a high-dimensional covariance matrix is converted to one of estimating low dimensional coefficients of the polynomial regression …
Estimating Spatial Autocorrelation With Sampled Network Data
Spatial autocorrelation is a parameter of importance for network data analysis. To estimate spatial autocorrelation, maximum likelihood has been popularly used. However, its rigorous implementation requires the whole network to be observed. This is practically infeasible if network size is huge (e.g., Facebook, Twitter, Weibo, WeChat, etc.). In that case, one has to rely on sampled network data to infer about spatial autocorrelation. By doing so,…
A Statistical Model for Social Network Labeling
We consider a social network from which one observes not only network structure (i.e., nodes and edges) but also a set of labels (or tags, keywords) for each node (or user). These labels are self-created and closely related to the user’s career status, life style, personal interests, and many others. Thus, they are of great interest for online marketing. To model their joint behavior with network structure, a complete data model is developed. The…
Testing the Diagonality of a Large Covariance Matrix in a Regression Setting
In multivariate analysis, the covariance matrix associated with a set of variables of interest (namely response variables) commonly contains valuable information about the dataset. When the dimension of response variables is considerably larger than the sample size, it is a nontrivial task to assess whether there are linear relationships between the variables. It is even more challenging to determine whether a set of explanatory variables can exp…
Estimating Mixture of Gaussian Processes by Kernel Smoothing
When the functional data are not homogeneous, e.g., there exist multiple classes of functional curves in the dataset, traditional estimation methods may fail. In this paper, we propose a new estimation procedure for the Mixture of Gaussian Processes, to incorporate both functional and inhomogeneous properties of the data. Our method can be viewed as a natural extension of high-dimensional normal mixtures. However, the key difference is that smoot…
Varying Naïve Bayes Models With Applications to Classification of Chinese Text Documents
Document classification is an area of great importance for which many classification methods have been developed. However, most of these methods cannot generate time-dependent classification rules. Thus, they are not the best choices for problems with time-varying structures. To address this problem, we propose a varying naïve Bayes model, which is a natural extension of the naïve Bayes model that allows for time-dependent classification rule. Th…
Feature Screening for Ultrahigh Dimensional Categorical Data With Applications
Ultrahigh dimensional data with both categorical responses and categorical covariates are frequently encountered in the analysis of big data, for which feature screening has become an indispensable statistical tool. We propose a Pearson chi-square based feature screening procedure for categorical response with ultrahigh dimensional categorical covariates. The proposed procedure can be directly applied for detection of important interaction effect…
Robust Regression Shrinkage and Consistent Variable Selection Through the LAD-Lasso
The least absolute deviation (LAD) regression is a useful method for robust regression, and the least absolute shrinkage and selection operator (lasso) is a popular choice for shrinkage estimation and variable selection. In this article we combine these two classical ideas together to produce LAD-lasso. Compared with the LAD regression, LAD-lasso can do parameter estimation and variable selection simultaneously. Compared with the traditional lass…
Correlation between Indian Ocean summer monsoon and North Atlantic climate during the Holocene
Sem obras proeminentes nesta página.
Correlation between Indian Ocean summer monsoon and North Atlantic climate during the Holocene
Robust Regression Shrinkage and Consistent Variable Selection Through the LAD-Lasso
The least absolute deviation (LAD) regression is a useful method for robust regression, and the least absolute shrinkage and selection operator (lasso) is a popular choice for shrinkage estimation and variable selection. In this article we combine these two classical ideas together to produce LAD-lasso. Compared with the LAD regression, LAD-lasso can do parameter estimation and variable selection simultaneously. Compared with the traditional lass…
Estimating Mixture of Gaussian Processes by Kernel Smoothing
When the functional data are not homogeneous, e.g., there exist multiple classes of functional curves in the dataset, traditional estimation methods may fail. In this paper, we propose a new estimation procedure for the Mixture of Gaussian Processes, to incorporate both functional and inhomogeneous properties of the data. Our method can be viewed as a natural extension of high-dimensional normal mixtures. However, the key difference is that smoot…
Varying Naïve Bayes Models With Applications to Classification of Chinese Text Documents
Document classification is an area of great importance for which many classification methods have been developed. However, most of these methods cannot generate time-dependent classification rules. Thus, they are not the best choices for problems with time-varying structures. To address this problem, we propose a varying naïve Bayes model, which is a natural extension of the naïve Bayes model that allows for time-dependent classification rule. Th…
Feature Screening for Ultrahigh Dimensional Categorical Data With Applications
Ultrahigh dimensional data with both categorical responses and categorical covariates are frequently encountered in the analysis of big data, for which feature screening has become an indispensable statistical tool. We propose a Pearson chi-square based feature screening procedure for categorical response with ultrahigh dimensional categorical covariates. The proposed procedure can be directly applied for detection of important interaction effect…
Testing the Diagonality of a Large Covariance Matrix in a Regression Setting
In multivariate analysis, the covariance matrix associated with a set of variables of interest (namely response variables) commonly contains valuable information about the dataset. When the dimension of response variables is considerably larger than the sample size, it is a nontrivial task to assess whether there are linear relationships between the variables. It is even more challenging to determine whether a set of explanatory variables can exp…
A Statistical Model for Social Network Labeling
We consider a social network from which one observes not only network structure (i.e., nodes and edges) but also a set of labels (or tags, keywords) for each node (or user). These labels are self-created and closely related to the user’s career status, life style, personal interests, and many others. Thus, they are of great interest for online marketing. To model their joint behavior with network structure, a complete data model is developed. The…
Estimating Spatial Autocorrelation With Sampled Network Data
Spatial autocorrelation is a parameter of importance for network data analysis. To estimate spatial autocorrelation, maximum likelihood has been popularly used. However, its rigorous implementation requires the whole network to be observed. This is practically infeasible if network size is huge (e.g., Facebook, Twitter, Weibo, WeChat, etc.). In that case, one has to rely on sampled network data to infer about spatial autocorrelation. By doing so,…
Covariance Matrix Estimation via Network Structure
In this article, we employ a regression formulation to estimate the high-dimensional covariance matrix for a given network structure. Using prior information contained in the network relationships, we model the covariance as a polynomial function of the symmetric adjacency matrix. Accordingly, the problem of estimating a high-dimensional covariance matrix is converted to one of estimating low dimensional coefficients of the polynomial regression …
Sequential Text-Term Selection in Vector Space Models
Text mining has recently attracted a great deal of attention with the accumulation of text documents in all fields. In this article, we focus on the use of textual information to explain continuous variables in the framework of linear regressions. To handle the unstructured texts, one common practice is to structuralize the text documents via vector space models. However, using words or phrases as the basic analysis terms in vector space models i…
Interactive Geological Data Visualization in an Immersive Environment
Underground flow paths (UFP) often play an important role in the illustration of geological data by geologists, especially in illustrating geological data and revealing stratigraphic structures, which can help domain experts in their exploration of petroleum information. In this paper, we present a new immersive visualization tool to help domain experts better illustrate stratigraphic data. We use a visualization method based on bit-array-based 3…
Feature Screening for Massive Data Analysis by Subsampling
Modern statistical analysis often encounters massive datasets with ultrahigh-dimensional features. In this work, we develop a subsampling approach for feature screening with massive datasets. The approach is implemented by repeated subsampling of massive data and can be used for analyzing tasks with memory constraints. To conduct the procedure, we first calculate an R-squared screening measure (and related sample moments) based on subsamples. Sec…
A Note on Distributed Quantile Regression by Pilot Sampling and One-Step Updating
Quantile regression is a method of fundamental importance. How to efficiently conduct quantile regression for a large dataset on a distributed system is of great importance. We show that the popularly used one-shot estimation is statistically inefficient if data are not randomly distributed across different workers. To fix the problem, a novel one-step estimation method is developed with the following nice properties. First, the algorithm is comm…
Autoregressive Model With Spatial Dependence and Missing Data
We study herein an autoregressive model with spatially correlated error terms and missing data. A logistic regression model with completely observed covariates is used to model the missingness mechanism. An autoregressive model is used to accommodate time series dependence, and a spatial error model is used to capture spatial dependence. To estimate the model, a weighted least squares estimator is developed for the temporal component, and a weigh…
Learning Human Activity Patterns Using Clustered Point Processes With Active and Inactive States
Modeling event patterns is a central task in a wide range of disciplines. In applications such as studying human activity patterns, events often arrive clustered with sporadic and long periods of inactivity. Such heterogeneity in event patterns poses challenges for existing point process models. In this article, we propose a new class of clustered point processes that alternate between active and inactive states. The proposed model is flexible, h…
Network Gradient Descent Algorithm for Decentralized Federated Learning
We study a fully decentralized federated learning algorithm, which is a novel gradient descent algorithm executed on a communication-based network. For convenience, we refer to it as a network gradient descent (NGD) method. In the NGD method, only statistics (e.g., parameter estimates) need to be communicated, minimizing the risk of privacy. Meanwhile, different clients communicate with each other directly according to a carefully designed networ…
Optimal Subsampling Bootstrap for Massive Data
The bootstrap is a widely used procedure for statistical inference because of its simplicity and attractive statistical properties. However, the vanilla version of bootstrap is no longer feasible computationally for many modern massive datasets due to the need to repeatedly resample the entire data. Therefore, several improvements to the bootstrap method have been made in recent years, which assess the quality of estimators by subsampling the ful…
Academic literature recommendation in large-scale citation networks enhanced by large language models
Penalized Sparse Covariance Regression with High Dimensional Covariates
Covariance regression offers an effective way to model the large covariance matrix with the auxiliary similarity matrices. In this work, we propose a sparse covariance regression (SCR) approach to handle the potentially high-dimensional predictors (i.e., similarity matrices). Specifically, we use the penalization method to identify the informative predictors and estimate their associated coefficients simultaneously. We first investigate the Lasso…
Computer Science (15 obras) · Mathematics (15 obras) · Statistics (13 obras) · Artificial Intelligence (12 obras) · Estimator (9 obras) · Statistical Methods and Inference (9 obras) · Data mining (8 obras) · Econometrics (6 obras) · Covariance (5 obras) · Advanced Statistical Methods and Models (4 obras)