|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | 基于动态LDA的科研文献主题演化分析显示文摘准确把握科研领域内文献主题的演化情况,有助于更好的进行科学研究。针对文献语料具有的单一主题时间性强、多个主题间关联性大等特点,本文在标准LDA模型基础上,将语料按照时序关系进行分片,建立动态LDA模型,以此来研究各个主题的强度和内容随时间的变化情况。同时选择目前热门的'Big Data'技术作为实验对象,从Web of Science数据库中抽取引文信息建立训练数据集,利用变分贝叶斯推断法对模型进行了求解,并对结果进行了可视化展示,实验表明,该方法简单有效,可以为把握科研发展趋势提供有效决策支持。 | 曾利 李自力 谭跃进 | 2014 | 软件2014,35,5: | 14 |
| 2 | A view on big data and its relation to Informetrics显示文摘Purpose:Big data offer a huge challenge.Their very existence leads to the contradiction that the more data we have the less accessible they become,as the particular piece of information one is searching for may be buried among terabytes of other data.In this contribution we discuss the origin of big data and point to three challenges when big data arise:Data storage,data processing and generating insights.Design/methodology/approach:Computer-related challenges can be expressed by the CAP theorem which states that it is only possible to simultaneously provide any two of the three following properties in distributed applications:Consistency(C),availability(A)and partition tolerance(P).As an aside we mention Amdahl’s law and its application for scientific collaboration.We further discuss data mining in large databases and knowledge representation for handling the results of data mining exercises.We further offer a short informetric study of the field of big data,and point to the ethical dimension of the big data phenomenon.Findings:There still are serious problems to overcome before the field of big data can deliver on its promises.Implications and limitations:This contribution offers a personal view,focusing on the information science aspects,but much more can be said about software aspects.Originality/value:We express the hope that the information scientists,including librarians,will be able to play their full role within the knowledge discovery,data mining and big data communities,leading to exciting developments,the reduction of scientific bottlenecks and really innovative applications. | Ronald ROUSSEAU | 2012 | Chinese Journal of Library and Information Science2012,,3: | 13 |
| 3 | 基于电信运营商的大数据解决方案分析显示文摘移动互联网时代,云计算、物联网、智能终端等新技术新应用涌现,移动互联网的迅猛发展给电信运营商带来了流量收益的同时,也带来了新的机遇和技术挑战。本文主要分析了大数据技术体系及其行业实践,大数据解决方案及典型产品的优缺点,并提出了一种适合电信运营商的大数据平台架构。 | 冯明丽 陈志彬 | 2013 | 通信与信息技术2013,,5: | 11 |
| 4 | Foundation Study on Wireless Big Data: Concept, Mining, Learning and Practices显示文摘Facing the development of future 5 G, the emerging technologies such as Internet of things, big data, cloud computing, and artificial intelligence is enhancing an explosive growth in data traffic. Radical changes in communication theory and implement technologies, the wireless communications and wireless networks have entered a new era. Among them, wireless big data(WBD) has tremendous value, and artificial intelligence(AI) gives unthinkable possibilities. However, in the big data development and artificial intelligence application groups, the lack of a sound theoretical foundation and mathematical methods is regarded as a real challenge that needs to be solved. From the basic problem of wireless communication, the interrelationship of demand, environment and ability, this paper intends to investigate the concept and data model of WBD, the wireless data mining, the wireless knowledge and wireless knowledge learning(WKL), and typical practices examples, to facilitate and open up more opportunities of WBD research and developments. Such research is beneficial for creating new theoretical foundation and emerging technologies of future wireless communications. | Jinkang Zhu Chen Gong Sihai Zhang Ming Zhao Wuyang Zhou | 2018 | China Communications2018,15,12: | 9 |
| 5 | 面向Big Data的数据处理技术概述显示文摘无所不在的移动设备、RFID、无线传感器每分每秒都在产生数据,数以亿计用户的互联网服务时时刻刻在产生巨量的交互。Big Data作为一个专有名词成为热点,归功于近年来互联网、云计算、移动和物联网的迅猛发展。针对现阶段业务需求和竞争压力对Big Data处理的实时性、有效性的高要求,本文在介绍面向Big Data处理方面的主要问题和难点的基础上,将现有的各种方法概括为两类并分别进行了阐述和分析,最后指出了该领域可能的发展方向。 | 夏海元 | 2012 | 数字技术与应用2012,30,3: | 9 |
| 6 | 基于MATLAB的音乐旋律二维可视化方法显示文摘音频可视化是信息可视化的重要分支,音乐是最具大众性和普遍性的音频信息,乐谱描述音乐的特点是专业性强,形式抽象。为了有利于音乐的展示,提出将音乐进行二维图形映射的可视化处理方法。旋律是音乐的基本要素,主要包含音高、时值和响度等特征,将音乐的旋律进行数据化,绘制二维图形,可以增强人们对音乐的感觉和认知。音频数据受到采样频率和分割帧的影响,会产生大量的过程性数据。MATLAB提供音频处理函数和大规模数据的分布式并行处理功能,可以完成音乐旋律二维可视化的实时处理。利用傅里叶变换提取音乐旋律的基本特征,形成音高、时值和响度等音乐特征向量矩阵,编制程序完成旋律二维可视化图形的自动绘制。 | 张岩 吕梦儒 | 2018 | 沈阳师范大学学报(自然科学版)2018,36,4: | 9 |
| 7 | 大数据及其商业价值显示文摘随着互联网行业的快速发展,大数据(Big Data)被认为将会是IT产业中最热门、最具发展性的领域。本文主要分析了大数据的特点、蕴涵的商业价值和面临的问题,为电子商务、电信运营商等行业的数据应用发展提供了新的视角。 | 陈勇 | 2013 | 通信与信息技术2013,,1: | 8 |
| 8 | Efficient visualization techniques for high resolution remotely sensed data in a network environment显示文摘There are three major research hotspots in efficient visualization techniques of high resolution remotely sensed data in network environment: the data organiza-tion and access in disk storage,the image data stitching and fitting methods,and the network transfers and access. In this paper a new method of 'Big File' organi-zation for improving the storage access efficiency of high resolution remote data is presented; a 'virtual data source' concept is introduced to solve the stitching problem of remotely sensed data from different sources with different resolutions; a remotely sensed data access engine design based on ATL technique is discussed to process the network transfers and access of remotely sensed data. All these techniques have been adopted in a prototype of digital China named 'ChinaStar'. | WENG JingNong1,WU Lun2,HUANG Jian1,XIA YuBin3 & CAI Heng1 1 College of Software,Beihang University,Beijing 100083,China 2 Institute of Geographical Information System and Remote Sensing,Peking University,Beijing 100871,China 3 College of Computer,Beihang University,Beijing 100083,China | 2008 | Science China(Technological Sciences)2008,51,S1: | 8 |
| 9 | Real-time intelligent big data processing:technology, platform, and applications显示文摘Human beings keep exploring the physical space using information means. Only recently, with the rapid development of information technologies and the increasing accumulation of data, human beings can learn more about the unknown world with data-driven methods. Given data timeliness, there is a growing awareness of the importance of real-time data. There are two categories of technologies accounting for data processing: batching big data and streaming processing, which have not been integrated well. Thus,we propose an innovative incremental processing technology named after Stream Cube to process both big data and stream data. Also, we implement a real-time intelligent data processing system, which is based on real-time acquisition, real-time processing, real-time analysis, and real-time decision-making. The real-time intelligent data processing technology system is equipped with a batching big data platform, data analysis tools, and machine learning models. Based on our applications and analysis, the real-time intelligent data processing system is a crucial solution to the problems of the national society and economy. | Tongya ZHENG Gang CHEN Xinyu WANG Chun CHEN Xingen WANG Sihui LUO | 2019 | Science China(Information Sciences)2019,62,8: | 8 |
| 10 | Research Progress of the Application of Big Data in China's Urban Planning显示文摘The arrival of the big data era facilitates the reform on related research and its application in the fi eld of urban planning from the mode of thinking to the technical method, which provides a technical foundation and platform for supporting the demands of residents as micro subjects in the process of city development. Beginning with analyzing the changes in urban planning thoughts under the infl uence of big data, this paper then summarizes the major reforms in urban planning research methodology and a city's plan formulation process caused by the application of big data, such as thinking mode, researching method, and planning process, based on which relevant issues like publicity and sharing of data, authenticity of data, and security of data that need to be further discussed and solved are analyzed, with the expectation of promoting the development of urban planning research and application in this new era. | Dang Anrong Xu Jian Tong Biao Li Juan Qian Fang | 2015 | China City Planning Review2015,24,1: | 7 |
| 11 | Efficient Bayesian networks for slope safety evaluation with large quantity monitoring information显示文摘New sensing and wireless technologies generate massive data. This paper proposes an efficient Bayesian network to evaluate the slope safety using large-quantity field monitoring information with underlying physical mechanisms. A Bayesian network for a slope involving correlated material properties and dozens of observational points is constructed. | Xueyou Li Limin Zhang Shuai Zhang | 2018 | Geoscience Frontiers2018,9,6: | 7 |
| 12 | The Application of Big data Mining in Risk Warning for Food Safety显示文摘Comprehensive evaluation and warning is very important and difficult in food safety. This paper mainly focuses on introducing the application of using big data mining in food safety warning field. At first,we introduce the concept of big data miming and three big data methods. At the same time,we discuss the application of the three big data miming methods in food safety areas. Then we compare these big data miming methods,and propose how to apply Back Propagation Neural Network in food safety risk warning. | Yajie WANG Bing YANG Yan LUO Jinlin HE Hong TAN | 2015 | Asian Agricultural Research2015,7,8: | 6 |
| 13 | Quality control of marine big data——a case study of real-time observation station data in Qingdao显示文摘Offshore waters provide resources for human beings,while on the other hand,threaten them because of marine disasters.Ocean stations are part of offshore observation networks,and the quality of their data is of great significance for exploiting and protecting the ocean.We used hourly mean wave height,temperature,and pressure real-time observation data taken in the Xiaomaidao station(in Qingdao,China)from June 1,2017,to May 31,2018,to explore the data quality using eight quality control methods,and to discriminate the most effective method for Xiaomaidao station.After using the eight quality control methods,the percentages of the mean wave height,temperature,and pressure data that passed the tests were 89.6%,88.3%,and 98.6%,respectively.With the marine disaster(wave alarm report)data,the values failed in the test mainly due to the influence of aging observation equipment and missing data transmissions.The mean wave height is often affected by dynamic marine disasters,so the continuity test method is not effective.The correlation test with other related parameters would be more useful for the mean wave height. | QIAN Chengcheng LIU Aichao HUANG Rui LIU Qingrong XU Wenkun ZHONG Shan YU Le | 2019 | Journal of Oceanology and Limnology2019,37,6: | 6 |
| 14 | A review of control loop monitoring and diagnosis:Prospects of controller maintenance in big data era显示文摘Owing to wide applications of automatic control systems in the process industries, the impacts of controller performance on industrial processes are becoming increasingly significant. Consequently, controller maintenance is critical to guarantee routine operations of industrial processes. The workflow of controller maintenance generally involves the following steps: monitor operating controller performance and detect performance degradation, diagnose probable root causes of control system malfunctions, and take specific actions to resolve associated problems. In this article, a comprehensive overview of the mainstream of control loop monitoring and diagnosis is provided, and some existing problems are also analyzed and discussed. From the viewpoint of synthesizing abundant information in the context of big data, some prospective ideas and promising methods are outlined to potentially solve problems in industrial applications. | Xinqing Gao Fan Yang Chao Shang Dexian Huang | 2016 | Chinese Journal of Chemical Engineering2016,24,8: | 6 |
| 15 | Big Data Analytics for Healthcare Industry:Impact,Applications,and Tools显示文摘In recent years, huge amounts of structured, unstructured, and semi-structured data have been generated by various institutions around the world and, collectively, this heterogeneous data is referred to as big data. The health industry sector has been confronted by the need to manage the big data being produced by various sources,which are well known for producing high volumes of heterogeneous data. Various big-data analytics tools and techniques have been developed for handling these massive amounts of data, in the healthcare sector. In this paper, we discuss the impact of big data in healthcare, and various tools available in the Hadoop ecosystem for handling it. We also explore the conceptual architecture of big data analytics for healthcare which involves the data gathering history of different branches, the genome database, electronic health records, text/imagery, and clinical decisions support system. | Sunil Kumar Maninder Singh | 2019 | Big Data Mining and Analytics2019,2,1: | 6 |
| 16 | Earth observation big data for climate change research显示文摘Earth observation technology has provided highly useful information in global climate change research over the past few decades and greatly promoted its development,especially through providing biological,physical,and chemical parameters on a global scale.Earth observation data has the 4V features(volume,variety,veracity,and velocity) of big data that are suitable for climate change research.Moreover,the large amount of data available from scientific satellites plays an important role.This study reviews the advances of climate change studies based on Earth observation big data and provides examples of case studies that utilize Earth observation big data in climate change research,such as synchronous satelliteeaerialeground observation experiments,which provide extremely large and abundant datasets; Earth observational sensitive factors(e.g.,glaciers,lakes,vegetation,radiation,and urbanization); and global environmental change information and simulation systems.With the era of global environment change dawning,Earth observation big data will underpin the Future Earth program with a huge volume of various types of data and will play an important role in academia and decisionmaking.Inevitably,Earth observation big data will encounter opportunities and challenges brought about by global climate change. | GUO Hua-Dong ZHANG Li ZHU Lan-Wei | 2015 | Advances in Climate Change Research2015,6,2: | 6 |
| 17 | Semi-supervised multi-layered clustering model for intrusion detection显示文摘A Machine Learning (ML)-based Intrusion Detection and Prevention System (IDPS)requires a large amount of labeled up-to-date training data to effectively detect intrusions and generalize well to novel attacks.However,the labeling of data is costly and becomes infeasible when dealing with big data,such as those generated by Intemet of Things applications.To this effect,building an ML model that learns from non-labeled or partially labeled data is of critical importance.This paper proposes a Semi-supervised Mniti-Layered Clustering ((SMLC))model for the detection and prevention of network intrusion.SMLC has the capability to learn from partially labeled data while achieving a detection performance comparable to that of supervised ML-based IDPS.The performance of SMLC is compared with that of a well-known semi-supervised model (tri-training)and of supervised ensemble ML models, namely Random.Forest,Bagging,and AdaboostM1on two benchmark network-intrusion datasets,NSL and Kyoto 2006+.Experimental resnits show that SMLC is superior to tri-training,providing a comparable detection accuracy with 20%less labeled instances of training data.Furthermore,our results demonstrate that our scheme has a detection accuracy comparable to that of the supervised ensemble models. | Omar Y.Al-Jarrah Yousof A1-Hammdi Patti D.Yoo Sami Muhaidat Mahmoud Al-Qutayri | 2018 | Digital Communications and Networks2018,4,4: | 6 |
| 18 | Comparative study of microarray and experimental data on Schwann cells in peripheral nerve degeneration and regeneration: big data analysis显示文摘A Schwann cell has regenerative capabilities and is an important cell in the peripheral nervous system.This microarray study is part of a bioinformatics study that focuses mainly on Schwann cells. Microarray data provide information on differences between microarray-based and experiment-based gene expression analyses. According to microarray data, several genes exhibit increased expression(fold change) but they are weakly expressed in experimental studies(based on morphology, protein and mRNA levels). In contrast, some genes are weakly expressed in microarray data and highly expressed in experimental studies;such genes may represent future target genes in Schwann cell studies. These studies allow us to learn about additional genes that could be used to achieve targeted results from experimental studies. In the current big data study by retrieving more than 5000 scientific articles from PubMed or NCBI, Google Scholar, and Google, 1016(up-and downregulated) genes were determined to be related to Schwann cells. However,no experiment was performed in the laboratory; rather, the present study is part of a big data analysis. Our study will contribute to our understanding of Schwann cell biology by aiding in the identification of genes.Based on a comparative analysis of all microarray data, we conclude that the microarray could be a good tool for predicting the expression and intensity of different genes of interest in actual experiments. | Ulfuara Shefa Junyang Jung | 2019 | Neural Regeneration Research2019,14,6: | 6 |
| 19 | A Novel Clustering Technique for Efficient Clustering of Big Data in Hadoop Ecosystem显示文摘Big data analytics and data mining are techniques used to analyze data and to extract hidden information.Traditional approaches to analysis and extraction do not work well for big data because this data is complex and of very high volume. A major data mining technique known as data clustering groups the data into clusters and makes it easy to extract information from these clusters. However, existing clustering algorithms, such as k-means and hierarchical, are not efficient as the quality of the clusters they produce is compromised. Therefore, there is a need to design an efficient and highly scalable clustering algorithm. In this paper, we put forward a new clustering algorithm called hybrid clustering in order to overcome the disadvantages of existing clustering algorithms. We compare the new hybrid algorithm with existing algorithms on the bases of precision, recall, F-measure, execution time, and accuracy of results. From the experimental results, it is clear that the proposed hybrid clustering algorithm is more accurate, and has better precision, recall, and F-measure values. | Sunil Kumar Maninder Singh | 2019 | Big Data Mining and Analytics2019,2,4: | 5 |
| 20 | Big Data Thinking and Its Biomedical Application显示文摘Big data thinking gradually rise with the coming era of big data. Big data characteristics could be summarized with 4V: volume, variety, velocity and value. The characteristics of big data thinking could be summed up in integrity, fault tolerance, correlation and intelligence. These characteristics were also the primary differences between big data thinking and small data thinking. The application of big data thinking in biomedical field became more and more widely, and the most commonly used was NCBI database. The process of mining valuable information in NCBI database was big data thinking. And the rise of Meta analysis and TCGA database would illustrate the huge application value of big data thinking in the biomedical field. | Petar Melih INAL Nikhil Vishnu | 2018 | Biomed Communication2018,2,1: | 5 |