|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | How do machine learning techniques help in increasing accuracy of landslide susceptibility maps?显示文摘Landslides are abundant in mountainous regions.They are responsible for substantial damages and losses in those areas.The A1 Highway,which is an important road in Algeria,was sometimes constructed in mountainous and/or semi-mountainous areas.Previous studies of landslide susceptibility mapping conducted near this road using statistical and expert methods have yielded ordinary results.In this research,we are interested in how do machine learning techniques help in increasing accuracy of landslide susceptibility maps in the vicinity of the A1 Highway corridor.To do this,an important section at Ain Bouziane(NE,Algeria) is chosen as a case study to evaluate the landslide susceptibility using three different machine learning methods,namely,random forest(RF),support vector machine(SVM),and boosted regression tree(BRT).First,an inventory map and nine input factors were prepared for landslide susceptibility mapping(LSM) analyses.The three models were constructed to find the most susceptible areas to this phenomenon.The results were assessed by calculating the receiver operating characteristic(ROC) curve,the standard error(Std.error),and the confidence interval(CI) at 95%.The RF model reached the highest predictive accuracy(AUC=97.2%) comparatively to the other models.The outcomes of this research proved that the obtained machine learning models had the ability to predict future landslide locations in this important road section.In addition,their application gives an improvement of the accuracy of LSMs near the road corridor.The machine learning models may become an important prediction tool that will identify landslide alleviation actions. | Yacine Achour Hamid Reza Pourghasemi | 2020 | Geoscience Frontiers2020,11,3: | 11 |
| 2 | 基于Stacking算法的员工离职预测分析与研究显示文摘针对员工离职会增加企业运营成本,降低企业盈利能力的问题,提出使用机器学习的离职员工预测算法;通过Stacking集成学习算法组合Adaboost和Random Forest基本算法构建LRA预测模型,实现对某企业的员工离职预测;实验结果显示,LRA模型的预测准确率为89. 09%,相对于单一算法所构建验证的模型预测准确率明显提高,LRA模型的查准率、查全率以及F1度量指标证实模型的可行性与可靠性,通过对输入LRA模型的特征进行重要性排序,得到影响员工离职的主要因素有加班、工龄(0-3年)、收入、职业级别等,丰富已有研究的结论,有利于企业决策者,针对离职行为进行合理决策。 | 李强 翟亮 | 2019 | 重庆工商大学学报(自然科学版)2019,36,1: | 9 |
| 3 | A Model Output Machine Learning Method for Grid Temperature Forecasts in the Beijing Area显示文摘In this paper, the model output machine learning (MOML) method is proposed for simulating weather consultation, which can improve the forecast results of numerical weather prediction (NWP). During weather consultation, the forecasters obtain the final results by combining the observations with the NWP results and giving opinions based on their experience. It is obvious that using a suitable post-processing algorithm for simulating weather consultation is an interesting and important topic. MOML is a post-processing method based on machine learning, which matches NWP forecasts against observations through a regression function. By adopting different feature engineering of datasets and training periods, the observational and model data can be processed into the corresponding training set and test set. The MOML regression function uses an existing machine learning algorithm with the processed dataset to revise the output of NWP models combined with the observations, so as to improve the results of weather forecasts. To test the new approach for grid temperature forecasts, the 2-m surface air temperature in the Beijing area from the ECMWF model is used. MOML with different feature engineering is compared against the ECMWF model and modified model output statistics (MOS) method. MOML shows a better numerical performance than the ECMWF model and MOS, especially for winter. The results of MOML with a linear algorithm, running training period, and dataset using spatial interpolation ideas, are better than others when the forecast time is within a few days. The results of MOML with the Random Forest algorithm, year-round training period, and dataset containing surrounding gridpoint information, are better when the forecast time is longer. | Haochen LI Chen YU Jiangjiang XIA Yingchun WANG Jiang ZHU Pingwen ZHANG | 2019 | Advances in Atmospheric Sciences2019,36,10: | 9 |
| 4 | Environmental factors influencing snowfall and snowfall prediction in the Tianshan Mountains, Northwest China显示文摘Snowfall is one of the dominant water resources in the mountainous regions and is closely related to the development of the local ecosystem and economy. Snowfall predication plays a critical role in understanding hydrological processes and forecasting natural disasters in the Tianshan Mountains, where meteorological stations are limited. Based on climatic, geographical and topographic variables at 27 meteorological stations during the cold season(October to April) from 1980 to 2015 in the Tianshan Mountains located in Xinjiang of Northwest China, we explored the potential influence of these variables on snowfall and predicted snowfall using two methods: multiple linear regression(MLR) model(a conventional measuring method) and random forest(RF) model(a non-parametric and non-linear machine learning algorithm). We identified the primary influencing factors of snowfall by ranking the importance of eight selected predictor variables based on the relative contribution of each variable in the two models. Model simulations were compared using different performance indices and the results showed that the RF model performed better than the MLR model, with a much higher R^2 value(R^2=0.74; R^2, coefficient of determination) and a lower bias error(RSR=0.51; RSR, the ratio of root mean square error to standard deviation of observed dataset). This indicates that the non-linear trend is more applicable for explaining the relationship between the selected predictor variables and snowfall. Relative humidity, temperature and longitude were identified as three of the most important variables influencing snowfall and snowfall prediction in both models, while elevation, aspect and latitude were of secondary importance, followed by slope and wind speed. These results will be beneficial to understand hydrological modeling and improve management and prediction of water resources in the Tianshan Mountains. | ZHANG Xueting LI Xuemei LI Lanhai ZHANG Shan QIN Qirui | 2019 | Journal of Arid Land2019,11,1: | 8 |
| 5 | Mapping landslide susceptibility at the Three Gorges Reservoir, China, using gradient boosting decision tree,random forest and information value models显示文摘This work was to generate landslide susceptibility maps for the Three Gorges Reservoir(TGR) area, China by using different machine learning models. Three advanced machine learning methods, namely, gradient boosting decision tree(GBDT), random forest(RF) and information value(InV) models, were used, and the performances were assessed and compared. In total, 202 landslides were mapped by using a series of field surveys, aerial photographs, and reviews of historical and bibliographical data. Nine causative factors were then considered in landslide susceptibility map generation by using the GBDT, RF and InV models. All of the maps of the causative factors were resampled to a resolution of 28.5 m. Of the 486289 pixels in the area,28526 pixels were landslide pixels, and 457763 pixels were non-landslide pixels. Finally, landslide susceptibility maps were generated by using the three machine learning models, and their performances were assessed through receiver operating characteristic(ROC) curves, the sensitivity, specificity,overall accuracy(OA), and kappa coefficient(KAPPA). The results showed that the GBDT, RF and In V models in overall produced reasonable accurate landslide susceptibility maps. Among these three methods, the GBDT method outperforms the other two machine learning methods, which can provide strong technical support for producing landslide susceptibility maps in TGR. | CHEN Tao ZHU Li NIU Rui-qing TRINDER C John PENG Ling LEI Tao | 2020 | Journal of Mountain Science2020,17,3: | 7 |
| 6 | Spatiotemporal evolution of carbon sequestration of limestone weathering in China显示文摘Carbonate carbon sequestration(CS) can aid in solving the problem of terrestrial residual carbon sinks and imbalances in the global carbon budget. Thus, complete understanding of the magnitude, spatiotemporal distribution, and evolution of this sequestration is highly desirable. On the basis of random forest regression and maximal potential dissolution model for carbonate, we estimated the CS of typical carbonate weathering in China from 2000 to 2014, that is, the sequestration of limestone weathering, using long-term ecologic, meteorological, hydrological raster data, and monitored data from 44 watersheds in China and surrounding regions. We extended our analyses by systematically exploring the spatiotemporal pattern and evolution trend of the flux and total sequestration. High levels of ionic activity coefficients of Ca^(2+) and HCO_3^- in limestone regions were observed to be mainly distributed in Northern and Northwestern China with a clear gradient from northwest to southeast. With a contrary spatial pattern, the annual average CS flux(CSF) of limestone weathering in China was estimated to be 4.28 t C km^(-2) yr^(-1), with high values mainly in the karst zones in Southeastern China. The mean CSF in different latitudes showed that Southern China(south of 28.14°N) was the region with the largest interannual fluctuation of flux and CSF increases as latitude decreases. The mean CSF in subtropical and tropical(TR) regions was the maximum of all major climate types, and for the frigid(F), mid-temperate(MTE), warm temperate(WTE), and temperate(TE) major climates; the CSF in the desert(D)subdivided climate was the minimum of these climates. By contrast, the values in grassland(G) and broad-leaved forest subdivided climate were the maximum. The pixel-based trend analysis indicated that the CSF of limestone weathering in China was slightly increasing in the period 2000–2014 with a rate of 0.036 t C km^(-2) yr^(-1). Furthermore, the annual total CS was estimated to be 7.07 Tg carbon per year(Tg C yr^(-1)) with high levels in 2002, 2008, and 2010, and the minimum appeared in 2011 with a slightly increasing trend of the total CS being observed with a rate of 0.06 Tg C yr^(-1). Tibet Autonomous Region was the administrative division with the largest total CS of limestone weathering(1.20 Tg C yr^(-1)) in China, and karst zones in Southeastern China had the largest total CS(4.95 Tg C yr^(-1)) which accounts for 70.01% of that in the three divided karst regions. On the basis of the diversity of rock chemical weathering carbon cycle mechanisms of different carbonate rock types, we estimated that the total CS of carbonate weathering in China may reach 11.37 Tg C yr^(-1)(the sink was approximately 5.02 t C km^(-2) yr^(-1)),which amounts to 16.20% of the total biomass CS in China, furthermore, the CSF of carbonate weathering in China can reach6.54 t C km^(-2) yr^(-1) if excluding the interference of the negative runoff. This finding indicates that CS of carbonate weathering is an indispensable part of China's terrestrial carbon sink system. The research pattern of this study was important for further improving the accuracy of the estimation for the global carbonate weathering carbon sink. | Huiwen LI Shijie WANG Xiaoyong BAI Yue CAO Luhua WU | 2019 | Science China Earth Sciences2019,62,6: | 7 |
| 7 | 1980-2017年中国碳强度关键影响因子变化特征显示文摘The Chinese government ratified the Paris Climate Agreement in 2016.Accordingly,China aims to reduce carbon dioxide emissions per unit of gross domestic product(carbon intensity)to 60%–65%of 2005 levels by 2030.However,since numerous factors influence carbon intensity in China,it is critical to assess their relative importance to determine the most important factors.As traditional methods are inadequate for identifying key factors from a range of factors acting in concert,machine learning was applied in this study.Specifically,random forest algorithm,which is based on decision tree theory,was employed because it is insensitive to multicollinearity,is robust to missing and unbalanced data,and provides reasonable predictive results.We identified the key factors affecting carbon intensity in China using random forest algorithm and analyzed the evolution in the key factors from 1980 to 2017.The dominant factors affecting carbon intensity in China from 1980 to 1991 included the scale and proportion of energy-intensive industry,the proportion of fossil fuel-based energy,and technological progress.The Chinese economy developed rapidly between 1992 and 2007;during this time,the effects of the proportion of service industry,price of fossil fuel,and traditional residential consumption on carbon intensity increased.Subsequently,the Chinese economy entered a period of structural adjustment after the 2008 global financial crisis;during this period,reductions in emissions and the availability of new energy types began to have effects on carbon intensity,and the importance of residential consumption increased.The results suggest that optimizing the energy and industrial structures,promoting technological advancement,increasing green consumption,and reducing emissions are keys to decreasing carbon intensity within China in the future.These approaches will help achieve the goal of reducing carbon intensity to 60%–65%of the 2005 level by 2030. | 唐志鹏 梅子傲 刘卫东 夏炎 | 2020 | Journal of Geographical Sciences2020,30,5: | 6 |
| 8 | 基于多种机器学习算法的波士顿房价预测显示文摘波士顿房价预测问题在人工智能领域中属于回归问题,回归问题是机器学习中很重要的一个研究方向。解决回归问题,较为常见的算法模型有Ridge Regression模型,基于集成学习方法的Random Forest、AdaBoost模型,基于深度学习的DNN模型等等。在具体问题中,选取的模型不同相应的效果也会不同,本研究根据'波士顿房价预测'这一实际问题,分别采用了不同的机器学习模型进行训练和测试,从多个角度对比了不同算法模型在房价预测这一回归问题上的综合表现,给出分析与总结。对上述不同算法模型的算法原理、实际表现以及特点进行了深入分析,总结了不同算法模型在波士顿房价预测这一实际问题上的表现,对不同模型的优缺点进行横向对比,对效果差异进行了分析与总结。 | 田润泽 | 2019 | 中国新通信2019,21,11: | 6 |
| 9 | Analysis of PM2.5 concentrations in Heilongjiang Province associated with forest cover and other factors显示文摘Atmospheric particulate matter(PM2.5) seriously influences air quality. It is considered one of the main environmental triggers for lung and heart diseases. Air pollutants can be adsorbed by forest. In this study we investigated the effect of forest cover on urban PM2.5 concentrations in 12 cities in Heilongjiang Province,China. The forest cover in each city was constant throughout the study period. The average daily concentration of PM2.5 in 12 cities was below 75 lg/m^3 during the non-heating period but exceeded this level during heating period. Furthermore, there were more moderate pollution days in six cities. This indicated that forests had the ability to reduce the concentration of PM2.5 but the main cause of air pollution was excessive human interference and artificial heating in winter. We classified the 12 cities according to the average PM2.5 concentrations. The relationship between PM2.5 concentrations and forest cover was obtained by integrating forest cover, land area,heated areas and number of vehicles in cities. Finally,considering the complexity of PM2.5 formation and based on the theory of random forestry, we selected six cities and analyzed their meteorological and air pollutant data. The main factors affecting PM2.5 concentrations were PM10,NO_2, CO and SO_2 in air pollutants while meteorological factors were secondary. | Yu Zheng San Li Chuanshan Zou Xiaojian Ma Guocai Zhang | 2019 | Journal of Forestry Research2019,30,1: | 6 |
| 10 | Integration of A Deep Learning Classifier with A Random Forest Approach for Predicting Malonylation Sites显示文摘As a newly-identified protein post-translational modification, malonylation is involved in a variety of biological functions. Recognizing malonylation sites in substrates represents an initial but crucial step in elucidating the molecular mechanisms underlying protein malonylation. In this study, we constructed a deep learning(DL) network classifier based on long short-term memory(LSTM) with word embedding(LSTMWE) for the prediction of mammalian malonylation sites.LSTMWEperforms better than traditional classifiers developed with common pre-defined feature encodings or a DL classifier based on LSTM with a one-hot vector. The performance of LSTMWE is sensitive to the size of the training set, but this limitation can be overcome by integration with a traditional machine learning(ML) classifier. Accordingly, an integrated approach called LEMP was developed, which includes LSTMWEand the random forest classifier with a novel encoding of enhanced amino acid content. LEMP performs not only better than the individual classifiers but also superior to the currently-available malonylation predictors. Additionally, it demonstrates a promising performance with a low false positive rate, which is highly useful in the prediction application. Overall, LEMP is a useful tool for easily identifying malonylation sites with high confidence.LEMP is available at http://www.bioinfogo.org/lemp. | Zhen Chen Ningning He Yu Huang Wen Tao Qin Xuhan Liu Lei Li | 2018 | Genomics, Proteomics & Bioinformatics2018,16,6: | 4 |
| 11 | Analyses of egg size,otolith shape,and growth revealed two components of small yellow croaker in Haizhou Bay spawning stock显示文摘The geographical variations in life history characteristics of small yellow croaker Larimichthys polyactis, caused by experienced different environmental conditions, have been observed in China seas. Previous studies based on spatial distribution, migration route, and body morphometrics suggested a complex stock structure. In this study, to clarify the source of a spawning stock, we investigated the reproduction strategy and inter-structure of the Haizhou Bay (HZB) spawning stock in the middle Yellow Sea from both egg survey and adult otolith increment analysis. Egg and adult samples were collected from three surveys during spawning season in 2013. Distinct spatial and temporal variations were detected in egg distribution and size, and otolith shape analysis of adult fishes revealed two morphotypes with different increment growth using random forest cluster. The results indicate the existence of two components within the same spawning stock in HZB from different wintering grounds, and accordingly special protection should be required for this stock given the significance to maintain connectivity between adjacent subpopulations. | JIANG Yiqian ZHANG Chi YE Zhenjiang TIAN Yongjun | 2019 | Journal of Oceanology and Limnology2019,37,4: | 3 |
| 12 | 改进随机森林算法在Android恶意软件检测中的应用显示文摘Random Forest作为一种常见的机器学习算法,不仅具备较高的分类回归性能,而且快速高效.传统的Random Forest算法并未在决策树的生成和选择上做深入研究,在本文中笔者提出一种降序去冗的寻优方式对机器学习中监督学习算法Random Forest进行改进,在保证准确率的同时减少随机森林的冗余度,并应用于Android系统的恶意软件检测.经过五折交叉验证法验证,改进的Random Forest算法能够在较低的冗余度下保证较高的准确率,同时改进的算法准确率在与同条件下的原算法的准确率以及OOB模型下的准确率相差在1%以内,在与单模型分类算法KNN和集成式学习算法Adaboost M1的对比试验中改进的Random Forest算法要优于以上两者. | 吴非 吴向前 陈晓燕 | 2017 | 新疆大学学报(自然科学版)2017,34,3: | 3 |
| 13 | Modeling oblique load carrying capacity of batter pile groups using neural network,random forest regression and M5 model tree显示文摘M5 model tree,random forest regression(RF)and neural network(NN)based modelling approaches were used to predict oblique load carrying capacity of batter pile groups using 247 laboratory experiments with smooth and rough pile groups.Pile length(L),angle of oblique load(a),sand density(ρ),number of batter piles(B),and number of vertical piles(V)as input and oblique load(Q)as output was used.Results suggest improved performance by RF regression for both pile groups.M5 model tree provides simple linear relation which can be used for the prediction of oblique load for field data also.Model developed using RF regression approach with smooth pile group data was found to be in good agreement for rough piles data.NN based approach was found performing equally well with both smooth and rough piles.Sensitivity analysis using all three modelling approaches suggest angle of oblique load(a)and number of batter pile(B)affect the oblique load capacity for both smooth and rough pile groups. | Tanvi SINGH Mahesh PAL V.K.ARORA | 2019 | Frontiers of Structural and Civil Engineering2019,13,3: | 3 |
| 14 | Classifying DNA Methylation Imbalance Data in Cancer Risk Prediction Using SMOTE and Tomek Link Methods显示文摘 | Chao Liu Jia Wu Labrador Mirador Yang Song Weiyan Hou | 2018 | 国际计算机前沿大会会议论文集2018,,2: | 2 |
| 15 | Influencing factors and growth state classification of a natural Metasequoia population显示文摘By analyzing the importance of influencing factors and conducting a comparative study of the effects of different sorting algorithms, a new method is proposed that is suitable for classifying the growth state of a natural Metasequoia glyptostroboides Hu and W.C. Cheng population. We studied 2817 M. glyptostroboides trees over 100 years old and analyzed their growth state by measuring 15 factors from stumpage, site condition, and environmental data. The dimensionality of all factors were reduced using the random forest algorithm, and we classified the remaining factors using the following algorithms: random forest, back-propagation(BP) neural networks, and support vector machine(SVM). The applicability of each sorting algorithm was analyzed. When all the d factors are used for classification and modeling, the model's overall accuracy,kappa coefficient and test accuracy were 85.5%, 0.739 and 85.8%, respectively. By reducing the dimensionality of the factors using the random forest algorithm, 11 factors most strongly influenced the classifications of the growth state of the Metasequoia population: diameter at breast height,height, crown width, age from stumpage data; longitude,latitude, elevation, slope aspect, gradient and slope position from the site condition data; and the edge of the field from the environmental data. For classifying the Metasequoia population, the random forest algorithm has the highest overall accuracy at 87.2%, which is 3.4 and 2.3% higher than the BP neural networks and SVM algorithms,respectively. The SVM algorithm is superior to the random forest algorithm with respect to classifying the state of mortality. The combination of the random forest and SVM algorithms and their combined information can be used to classify and predict the growth state of this natural M.glyptostroboides population to provide a scientific basis for its effective protection. | Mu Liu Zhongke Feng Chenghui Ma Liyan Yang | 2019 | Journal of Forestry Research2019,30,1: | 1 |
| 16 | Spatial Prediction of Soil Salinity in a Semiarid Oasis: Environmental Sensitive Variable Selection and Model Comparison显示文摘Timely monitoring and early warning of soil salinity are crucial for saline soil management. Environmental variables are commonly used to build soil salinity prediction model. However, few researches have been done to summarize the environmental sensitive variables for soil electrical conductivity(EC) estimation systematically. Additionally, the performance of Multiple Linear Regression(MLR), Geographically Weighted Regression(GWR), and Random Forest regression(RFR) model, the representative of current main methods for soil EC prediction, has not been explored. Taking the north of Yinchuan plain irrigation oasis as the study area, the feasibility and potential of 64 environmental variables, extracted from the Landsat 8 remote sensed images in dry season and wet season, the digital elevation model, and other data, were assessed through the correlation analysis and the performance of MLR, GWR, and RFR model on soil salinity estimation was compared. The results showed that: 1) 10 of 15 imagery texture and spectral band reflectivity environmental variables extracted from Landsat 8 image in dry season were significantly correlated with soil EC, while only 3 of these indices extracted from Landsat 8 image in wet season have significant correlation with soil EC. Channel network base level, one of the terrain attributes, had the largest absolute correlation coefficient of 0.47 and all spatial location factors had significant correlation with soil EC. 2) Prediction accuracy of RFR model was slightly higher than that of the GWR model, while MLR model produced the largest error. 3) In general, the soil salinization level in the study area gradually increased from south to north. In conclusion, the remote sensed imagery scanned in dry season was more suitable for soil EC estimation, and topographic factors and spatial location also play a key role. This study can contribute to the research on model construction and variables selection for soil salinity estimation in arid and semiarid regions. | LI Zhen LI Yong XING An ZHUO Zhiqing ZHANG Shiwen ZHANG Yuanpei HUANG Yuanfang | 2019 | Chinese Geographical Science2019,29,5: | 1 |
| 17 | 基于SVM和Random Forest的高光谱遥感影像分类显示文摘针对传统高光谱分类方法的局限性,近年来发展的支持向量机(Support Vector Machine,SVM和多分类器学习对提高高光谱影像分类精度有着巨大推动作用。多分类器学习的基本思想是通过对分类器集合中分类器的选择与组合,获得比任何单一分类器更高的精度。本文提出用Random Forest对高光谱遥感影像进行分类,并与常用的多分类器的Bagging和AdaBoost算法进行时比,发现能够PandomForest能够获得更好的结果。最后对SVM和Random Forest方法进行对比,发现两者能够提高不同类别的分类精度。 | 潘屹峰 黄晶 | 2012 | 华东科技(学术版)2012,,10: | 1 |
| 18 | SPT based determination of undrained shear strength:Regression models and machine learning显示文摘The purpose of this study is the accurate prediction of undrained shear strength using Standard Penetration Test results and soil consistency indices,such as water content and Atterberg limits.With this study,along with the conventional methods of simple and multiple linear regression models,three machine learning algorithms,random forest,gradient boosting and stacked models,are developed for prediction of undrained shear strength.These models are employed on a relatively large data set from different projects around Turkey covering 230 observations.As an improvement over the available studies in literature,this study utilizes correct statistical analyses techniques on a relatively large database,such as using a train/test split on the data set to avoid overfitting of the developed models.Furthermore,the validity and consistency of the prediction results are ensured with the correct use of statistical measures like p-value and cross-validation which were missing in previous studies.To compare the performances of the models developed in this study with the prior ones existing in literature,all models were applied on the test data set and their performances are evaluated in terms of the resulting root mean squared error(RMSE)values and coefficient of determination(R^2).Accordingly,the models developed in this study demonstrate superior prediction capabilities compared to all of the prior studies.Moreover,to facilitate the use of machine learning algorithms for prediction purposes,entire source code prepared for this study and the collected data set are provided as supplements of this study. | Walid Khalid MBARAK Esma Nur CINICIOGLU Ozer CINICIOGLU | 2020 | Frontiers of Structural and Civil Engineering2020,14,1: | 1 |
| 19 | Modeling the relationship of diverse genomic signatures to gene expression levels with the regulation of long-range enhancer-promoter interactions显示文摘Enhancer-promoter (E-P) interaction is an essential component of cis-regulatory regulation for gene expression. However;to comprehensively study the gene expression with the regulation of long-range E-P interactions is a major challenge in the regulatory networks. As these types of gene expression are regulated by diverse genomic signatures, we presented a computational method to study the relationships between gene expression levels and diverse genomic signatures. In this paper;based on the datasets of long-range E-P interactions, we extracted feature parameters from multiple signatures (e.g., epigenetic marks, transcription factors) and used regression models to predict the gene expression levels. In our results, we found that the predicted expression values correlated well w让h the measured expression values in both the interacting and non-interacting sets, and the correlation values of the interacting set were higher than that of the corresponding non-interacting set in each cell line, which indicated that the distal enhancers would cooperate w让h diverse genomic signatures to facilitate the expression level of target genes. By comparing the important signature features for the gene expression levels between the interacting and non-interacting sets in the same cell line, we found that the important specific signatures affect the gene expression regulated by distal enhancers. Our research provided additional insights about the roles of diverse signatures in gene expression with the regulation of distal enhancers. | Zhen-Xing Feng Qian-Zhong Li Jian-fun Meng | 2019 | Biophysics Reports2019,5,3: | 1 |
| 20 | RANDOM FOREST FOR INTERMEDIATE DESCRIPTOR FUSION IN SHOT BOUNDARY DETECTION显示文摘Shot boundary detection is the fundamental part in many real applications as video retrieval and so on. This paper tackles the problem of video segment obtaining in complex movie videos. Firstly, intermediate descriptor is proposed to depict the variation of both abrupt and gradual change in shot boundaries, which is formed by distance vector on Local Binary Pattern(LBP), GIST(GIST) or their fusion. Instead of just using the adjacent frames distance, intermediate descriptor keeps the distances between current frame and consecutive frames. It comprehensively characterizes local temporal structure, which is especially important for gradual change. For the excellent ability for feature fusion in random forests, it is adopted here to verify the fusion effect of intermediate descriptor on LBP and GIST. The whole experiments are designed on the subset of TRECVid 2013 INS(INstance Search) task to verify the effectiveness of proposed intermediate descriptor and the fusion ability for random forest. Compared with static and adaptive thresholds approaches, the best performance can be achieved by post-fusion of intermediate descriptor on LBP and GIST. | Zhang Lei Chang Anqi Xiang Xuezhi | 2014 | Journal of Electronics(China)2014,31,5: | 1 |