|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | Science Mapping:A Systematic Review of the Literature显示文摘Purpose: We present a systematic review of the literature concerning major aspects of science mapping to serve two primary purposes: First, to demonstrate the use of a science mapping approach to perform the review so that researchers may apply the procedure to the review of a scientific domain of their own interest, and second, to identify major areas of research activities concerning science mapping, intellectual milestones in the development of key specialties, evolutionary stages of major specialties involved, and the dynamics of transitions from one specialty to another.Design/methodology/approach: We first introduce a theoretical framework of the evolution of a scientific specialty. Then we demonstrate a generic search strategy that can be used to construct a representative dataset of bibliographic records of a domain of research. Next, progressively synthesized co-citation networks are constructed and visualized to aid visual analytic studies of the domain's structural and dynamic patterns and trends. Finally, trajectories of citations made by particular types of authors and articles are presented to illustrate the predictive potential of the analytic approach.Findings: The evolution of the science mapping research involves the development of a number of interrelated specialties. Four major specialties are discussed in detail in terms of four evolutionary stages: conceptualization, tool construction, application, and codification. Underlying connections between major specialties are also explored. The predictive analysis demonstrates citations trajectories of potentially transformative contributions.Research limitations: The systematic review is primarily guided by citation patterns in the dataset retrieved from the literature. The scope of the data is limited by the source of the retrieval, i.e. the Web of Science, and the composite query used. An iterative query refinement is possible if one would like to improve the data quality, although the current approach serves our purpose adequately. More in-depth analyses of each specialty would be more revealing by incorporating additional methods such as citation context analysis and studies of other aspects of scholarly publications.Practical implications: The underlying analytic process of science mapping serves many practical needs, notably bibliometric mapping, knowledge domain visualization, and visualization of scientific literature. In order to master such a complex process of science mapping, researchers often need to develop a diverse set of skills and knowledge that may span multiple disciplines. The approach demonstrated in this article provides a generic method for conducting a systematic review.Originality/value: Incorporating the evolutionary stages of a specialty into the visual analytic study of a research domain is innovative. It provides a systematic methodology for researchers to achieve a good understanding of how scientific fields evolve, to recognize potentially insightful patterns from visually encoded signs, and to synthesize various information so as to capture the state of the art of the domain. | Chaomei Chen | 2017 | Journal of Data and Information Science2017,2,2: | 596 |
| 2 | Big Earth data:A new frontier in Earth and information sciences显示文摘Big data is a revolutionary innovation that has allowed the development of many new methods in scientific research.This new way of thinking has encouraged the pursuit of new discoveries.Big data occupies the strategic high ground in the era of knowledge economies and also constitutes a new national and global strategic resource.“Big Earth data”,derived from,but not limited to,Earth observation has macro-level capabilities that enable rapid and accurate monitoring of the Earth,and is becoming a new frontier contributing to the advancement of Earth science and significant scientific discoveries.Within the context of the development of big data,this paper analyzes the characteristics of scientific big data and recognizes its great potential for development,particularly with regard to the role that big Earth data can play in promoting the development of Earth science.On this basis,the paper outlines the Big Earth Data Science Engineering Project(CASEarth)of the Chinese Academy of Sciences Strategic Priority Research Program.Big data is at the forefront of the integration of geoscience,information science,and space science and technology,and it is expected that big Earth data will provide new prospects for the development of Earth science. | Huadong Guo | 2017 | Big Earth Data2017,1,1: | 49 |
| 3 | FAIR Principles:Interpretations and Implementation Considerations显示文摘The FAIR principles have been widely cited,endorsed and adopted by a broad range of stakeholders since their publication in 2016.By intention,the 15 FAIR guiding principles do not dictate specific technological implementations,but provide guidance for improving Findability,Accessibility,Interoperability and Reusability of digital resources.This has likely contributed to the broad adoption of the FAIR principles,because individual stakeholder communities can implement their own FAIR solutions.However,it has also resulted in inconsistent interpretations that carry the risk of leading to incompatible implementations.Thus,while the FAIR principles are formulated on a high level and may be interpreted and implemented in different ways,for true interoperability we need to support convergence in implementation choices that are widely accessible and(re)-usable.We introduce the concept of FAIR implementation considerations to assist accelerated global participation and convergence towards accessible,robust,widespread and consistent FAIR implementations.Any self-identified stakeholder community may either choose to reuse solutions from existing implementations,or when they spot a gap,accept the challenge to create the needed solution,which,ideally,can be used again by other communities in the future.Here,we provide interpretations and implementation considerations(choices and challenges)for each FAIR principle. | Annika Jacobsen Ricardo de Miranda Azevedo Nick Juty Dominique Batista Simon Coles Ronald Cornet Melanie Courtot Merce Crosas Michel Dumontier Chris T.Evelo Carole Goble Giancarlo Guizzardi Karsten Kryger Hansen Ali Hasnain Kristina Hettne Jaap Heringa Rob W.W.Hooft Melanie Imming Keith G.Jeffery Rajaram Kaliyaperumal Martijn GKersloot Christine R.Kirkpatrick Tobias Kuhn Ignasi Labastida Barbara Magagna PeterMcQuilton Natalie Meyers Annalisa Montesanti Mirjam van Reisen Philippe Rocca-Serra Robert Pergl Susanna-Assunta Sansone Luiz Olavo Bonino da Silva Santos Juliane Schneider George Strawn Mark Thompson Andra Waagmeester Tobias Weigel Mark D.Wilkinson Egon L.Willighagen Peter Wittenburg Marco Roos Barend Mons Erik Schultes | 2020 | Data Intelligence2020,2,1: | 26 |
| 4 | Multi-view Clustering: A Survey显示文摘In the big data era, the data are generated from different sources or observed from different views. These data are referred to as multi-view data. Unleashing the power of knowledge in multi-view data is very important in big data mining and analysis. This calls for advanced techniques that consider the diversity of different views,while fusing these data. Multi-view Clustering(MvC) has attracted increasing attention in recent years by aiming to exploit complementary and consensus information across multiple views. This paper summarizes a large number of multi-view clustering algorithms, provides a taxonomy according to the mechanisms and principles involved, and classifies these algorithms into five categories, namely, co-training style algorithms, multi-kernel learning, multiview graph clustering, multi-view subspace clustering, and multi-task multi-view clustering. Therein, multi-view graph clustering is further categorized as graph-based, network-based, and spectral-based methods. Multi-view subspace clustering is further divided into subspace learning-based, and non-negative matrix factorization-based methods. This paper does not only introduce the mechanisms for each category of methods, but also gives a few examples for how these techniques are used. In addition, it lists some publically available multi-view datasets.Overall, this paper serves as an introductory text and survey for multi-view clustering. | Yan Yang Hao Wang | 2018 | Big Data Mining and Analytics2018,1,2: | 20 |
| 5 | Masked Sentence Model Based on BERT for Move Recognition in Medical Scientific Abstracts显示文摘Purpose:Mo ve recognition in scientific abstracts is an NLP task of classifying sentences of the abstracts into different types of language units.To improve the performance of move recognition in scientific abstracts,a novel model of move recognition is proposed that outperforms the BERT-based method.Design/methodology/approach:Prevalent models based on BERT for sentence classification often classify sentences without considering the context of the sentences.In this paper,inspired by the BERT masked language model(MLM),we propose a novel model called the masked sentence model that integrates the content and contextual information of the sentences in move recognition.Experiments are conducted on the benchmark dataset PubMed 20K RCT in three steps.Then,we compare our model with HSLN-RNN,BERT-based and SciBERT using the same dataset.Findings:Compared with the BERT-based and SciBERT models,the F1 score of our model outperforms them by 4.96%and 4.34%,respectively,which shows the feasibility and effectiveness of the novel model and the result of our model comes closest to the state-of-theart results of HSLN-RNN at present.Research limitations:The sequential features of move labels are not considered,which might be one of the reasons why HSLN-RNN has better performance.Our model is restricted to dealing with biomedical English literature because we use a dataset from PubMed,which is a typical biomedical database,to fine-tune our model.Practical implications:The proposed model is better and simpler in identifying move structures in scientific abstracts and is worthy of text classification experiments for capturing contextual features of sentences.Originality/value:T he study proposes a masked sentence model based on BERT that considers the contextual features of the sentences in abstracts in a new way.The performance of this classification model is significantly improved by rebuilding the input layer without changing the structure of neural networks. | Gaihong Yu Zhixiong Zhang Huan Liu Liangping Ding | 2019 | Journal of Data and Information Science2019,4,4: | 14 |
| 6 | Mapping landslide susceptibility and types using Random Forest显示文摘Landslides are one of the most destructive natural hazards;they can drastically alter landscape morphology,destroy man-made struc-tures,and endanger people’s life.Landslide susceptibility maps(LSMs),which show the spatial likelihood of landslide occurrence,are crucial for environmental management,urban planning,and minimizing economic losses.To date,the majority of research into data mining LSM uses small-scale case studies focusing on a single type of landslide.This paper presents a data mining approach to producing LSM for a large,heterogeneous region that is susceptible tomultipletypesoflandslides.UsingacasestudyofPiedmont,Italy,a Random Forest algorithm is applied to produce both susceptibility maps and classification maps.These maps are combined to give a highly accurate(over 85%classification accuracy)LSM which con-tains a large amount of information and is easy to interpret.This novel method of mapping landslide susceptibility demonstrates the efficacy of Random Forest to produce highly accurate susceptibility maps for alargeheterogeneousregion withouttheneed formultiple susceptibility assessments. | Khaled Taalab Tao Cheng Yang Zhang | 2018 | Big Earth Data2018,2,2: | 13 |
| 7 | Big Data and Data Science:Opportunities and Challenges of iSchools显示文摘Due to the recent explosion of big data, our society has been rapidly going through digital transformation and entering a new world with numerous eye-opening developments. These new trends impact the society and future jobs, and thus student careers. At the heart of this digital transformation is data science, the discipline that makes sense of big data. With many rapidly emerging digital challenges ahead of us, this article discusses perspectives on iSchools' opportunities and suggestions in data science education. We argue that iSchools should empower their students with 'information computing' disciplines, which we define as the ability to solve problems and create values, information, and knowledge using tools in application domains. As specific approaches to enforcing information computing disciplines in data science education, we suggest the three foci of user-based, tool-based, and applicationbased. These three foci will serve to differentiate the data science education of iSchools from that of computer science or business schools. We present a layered Data Science Education Framework (DSEF) with building blocks that include the three pillars of data science (people, technology, and data), computational thinking, data-driven paradigms, and data science lifecycles. Data science courses built on the top of this framework should thus be executed with user-based, tool-based, and application-based approaches. This framework will help our students think about data science problems from the big picture perspective and foster appropriate problem-solving skills in conjunction with broad perspectives of data science lifecycles. We hope the DSEF discussed in this article will help fellow iSchools in their design of new data science curricula. | Il-Yeol Song Yongjun Zhu | 2017 | Journal of Data and Information Science2017,2,3: | 12 |
| 8 | Patent Citations Analysis and Its Value in Research Evaluation: A Review and a New Approach to Map Technology-relevant Research显示文摘Purpose:First,to review the state-of-the-art in patent citation analysis,particularly characteristics of patent citations to scientific literature(scientific non-patent references,SNPRs).Second,to present a novel mapping approach to identify technology-relevant research based on the papers cited by and referring to the SNPRs.Design/methodology/approach:In the review part we discuss the context of SNPRs such as the time lags between scientific achievements and inventions.Also patent-to-patent citation is addressed particularly because this type of patent citation analysis is a major element in the assessment of the economic value of patents.We also review the research on the role of universities and researchers in technological development,with important issues such as universities as sources of technological knowledge and inventor-author relations.We conclude the review part of this paper with an overview of recent research on mapping and network analysis of the science and technology interface and of technological progress in interaction with science.In the second part we apply new techniques for the direct visualization of the cited and citing relations of SNPRs,the mapping of the landscape around SNPRs by bibliographic coupling and co-citation analysis,and the mapping of the conceptual environment of SNPRs by keyword co-occurrence analysis.Findings:We discuss several properties of SNPRs.Only a small minority of publications covered by the Web of Science or Scopus are cited by patents,about 3%–4%.However,for publications based on university-industry collaboration the number of SNPRs is considerably higher,around 15%.The proposed mapping methodology based on a 'second order SNPR approach' enables a better assessment of the technological relevance of research.Research limitations:The main limitation is that a more advanced merging of patent and publication data,in particular unification of author and inventor names,in still a necessity.Practical implications:The proposed mapping methodology enables the creation of a database of technology-relevant papers(TRPs).In a bibliometric assessment the publications of research groups,research programs or institutes can be matched with the TRPs and thus the extent to which the work of groups,programs or institutes are relevant for technological development can be measured.Originality/value:The review part examines a wide range of findings in the research of patent citation analysis.The mapping approach to identify a broad range of technologyrelevant papers is novel and offers new opportunities in research evaluation practices. | Anthony F.J. van Raan | 2017 | Journal of Data and Information Science2017,2,1: | 12 |
| 9 | FAIR Data and Services in Biodiversity Science and Geoscience显示文摘We examine the intersection of the FAIR principles(Findable,Accessible,Interoperable and Reusable),the challenges and opportunities presented by the aggregation of widely distributed and heterogeneous data about biological and geological specimens,and the use of the Digital Object Architecture(DOA)data model and components as an approach to solving those challenges that offers adherence to the FAIR principles as an integral characteristic.This approach will be prototyped in the Distributed System of Scientific Collections(DiSSCo)project,the pan-European Research Infrastructure which aims to unify over 110 natural science collections across 21 countries.We take each of the FAIR principles,discuss them as requirements in the creation of a seamless virtual collection of bio/geo specimen data,and map those requirements to Digital Object components and facilities such as persistent identification,extended data typing,and the use of an additional level of abstraction to normalize existing heterogeneous data structures.The FAIR principles inform and motivate the work and the DO Architecture provides the technical vision to create the seamless virtual collection vitally needed to address scientific questions of societal importance. | Larry Lannom Dimitris Koureas Alex R.Hardisty | 2020 | Data Intelligence2020,2,1: | 12 |
| 10 | Big data drives the development of Earth science显示文摘Big data is now a popular topic,becoming increasingly well known around the world,and yet the concept of big data and its implications are still novel.To discuss big data,it is appropriate to first talk about what is really meant by this term,and so to begin the first article in the inaugural issue of Big Earth Data,let us look at how data has become big data and why that is important. | Huadong Guo | 2017 | Big Earth Data2017,1,1: | 11 |
| 11 | AMiner:Search and Mining of Academic Social Networks显示文摘AMiner is a novel online academic search and mining system,and it aims to provide a systematic modeling approach to help researchers and scientists gain a deeper understanding of the large and heterogeneous networks formed by authors,papers,conferences,journals and organizations.The system is subsequently able to extract researchers’profiles automatically from the Web and integrates them with published papers by a way of a process that first performs name disambiguation.Then a generative probabilistic model is devised to simultaneously model the different entities while providing a topic-level expertise search.In addition,AMiner offers a set of researcher-centered functions,including social influence analysis,relationship mining,collaboration recommendation,similarity analysis and community evolution.The system has been in operation since 2006 and has been accessed from more than 8 million independent IP addresses residing in more than 200 countries and regions. | Huaiyu Wan Yutao Zhang Jing Zhang Jie Tang | 2019 | Data Intelligence2019,1,1: | 11 |
| 12 | DEEPEYE: Link Prediction in Dynamic Networks Based on Non-negative Matrix Factorization显示文摘A Non-negative Matrix Factorization(NMF)-based method is proposed to solve the link prediction problem in dynamic graphs. The method learns latent features from the temporal and topological structure of a dynamic network and can obtain higher prediction results. We present novel iterative rules to construct matrix factors that carry important network features and prove the convergence and correctness of these algorithms. Finally, we demonstrate how latent NMF features can express network dynamics efficiently rather than by static representation,thereby yielding better performance. The amalgamation of time and structural information makes the method achieve prediction results that are more accurate. Empirical results on real-world networks show that the proposed algorithm can achieve higher accuracy prediction results in dynamic networks in comparison to other algorithms. | Nahla Mohamed Ahmed Ling Chen Yulong Wang Bin Li Yun Li Wei Liu | 2018 | Big Data Mining and Analytics2018,1,1: | 11 |
| 13 | Unique,Persistent,Resolvable:Identifiers as the Foundation of FAIR显示文摘The FAIR principles describe characteristics intended to support access to and reuse of digital artifacts in the scientific research ecosystem.Persistent,globally unique identifiers,resolvable on the Web,and associated with a set of additional descriptive metadata,are foundational to FAIR data.Here we describe some basic principles and exemplars for their design,use and orchestration with other system elements to achieve FAIRness for digital research objects. | Nick Juty Sarala M.Wimalaratne Stian Soiland-Reyes John Kunze Carole A.Goble Tim Clark | 2020 | Data Intelligence2020,2,1: | 11 |
| 14 | A Bibliometric Framework for Identifying'Princes'Who Wake up the'Sleeping Beauty'in Challenge-type Scientific Discoveries显示文摘Purpose:This paper develops and validates a bibliometric framework for identifying the 'princes'(PR) who wake up the 'sleeping beauty'(SB) in challenge-type scientific discoveries,so as to figure out the awakening mechanisms,and promote potentially valuable but not readily accepted innovative research.(A PR is a research study.)Design/methodology/approach:We propose that PR candidates must meet the following four criteria:(1) be published near the time when the SB began to attract a lot of citations;(2) be highly cited papers themselves;(3) receive a substantial number of co-citations with the SB;and(4) within the challenge-type discoveries which contradict established theories,the 'pulling effect' of the PR on the SB must be strong.We test the usefulness of the bibliometric framework through a case study of a key publication by the 2014 chemistry Nobel laureate Stefan W.Hell,who negated Ernst Abbe's diffraction limit theory,one of the most prominent paradigms in the natural sciences.Findings:The first-ranked candidate PR article identified by the bibliometric framework is in line with historical facts.An SB may need one or more PRs and even 'retinues' to be 'awakened.' Documents with potential awakening functionality tend to be published in prestigious multidisciplinary journals with higher impact and wider scope than the journals publishing SBs.Research limitations:The above framework is only applicable to transformative innovations,and the conclusions are drawn from the analysis of one typical SB and her awakening process.Therefore the generality of our work might be limited.Practical implications:Publications belonging to so-called transformative research,even when less frequently cited,should be given special attention as early as possible,because they may suddenly attract many citations after a period of sleep,as reflected in our case study.Originality/value:The definition of PR(s) as the first paper(s) that cited the SB article(selfciting excluded) has its limitations.Instead,the SB-PR co-citations should be given priority in current environment of scholarly communication.Since the 'premature' or 'transformative' breakthroughs in the challenge-type SB documents are either beyond the current knowledge domain,or violate established paradigms,people's psychological distance from the SB is larger than that from the PR,which explains why the annual citations of the PR are usually higher than those of the SB,especially prior to or during the SB's citation boom period. | Jian Du YishanWu | 2016 | Journal of Data and Information Science2016,1,1: | 9 |
| 15 | Relation Classification via Recurrent Neural Network with Attention and Tensor Layers显示文摘Relation classification is a crucial component in many Natural Language Processing(NLP) systems. In this paper, we propose a novel bidirectional recurrent neural network architecture(using Long Short-Term Memory,LSTM, cells) for relation classification, with an attention layer for organizing the context information on the word level and a tensor layer for detecting complex connections between two entities. The above two feature extraction operations are based on the LSTM networks and use their outputs. Our model allows end-to-end learning from the raw sentences in the dataset, without trimming or reconstructing them. Experiments on the SemEval-2010 Task 8dataset show that our model outperforms most state-of-the-art methods. | Runyan Zhang Fanrong Meng Yong Zhou Bing Liu | 2018 | Big Data Mining and Analytics2018,1,3: | 9 |
| 16 | CN-DBpedia2: An Extraction and Verification Framework for Enriching Chinese Encyclopedia Knowledge Base显示文摘Knowledge base plays an important role in machine understanding and has been widely used in various applications, such as search engine, recommendation system and question answering. However, most knowledge bases are incomplete, which can cause many downstream applications to perform poorly because they cannot find the corresponding facts in the knowledge bases. In this paper, we propose an extraction and verification framework to enrich the knowledge bases. Specifically, based on the existing knowledge base, we first extract new facts from the description texts of entities. But not all newly-formed facts can be added directly to the knowledge base because the errors might be involved by the extraction. Then we propose a novel crowd-sourcing based verification step to verify the candidate facts. Finally, we apply this framework to the existing knowledge base CN-DBpedia and construct a new version of knowledge base CN-DBpedia2, which additionally contains the high confidence facts extracted from the description texts of entities. | Bo Xu Jiaqing Liang Chenhao Xie Bin Liang Lihan Chen Yanghua Xiao | 2019 | Data Intelligence2019,1,3: | 9 |
| 17 | Generation of ready to use (RTU) products over China based on Landsat series data显示文摘Earth observation community has entered into the era of big data.Family of Landsat sensors have collected massive medium resolution satellite images,which are valuable for long-term land surface monitoring.In order to significantly reduce the magnitude of data processing for remote sensing data users,Landsat-based Ready to Use(RTU)products have been produced.Main RTU products,including orthorectified products,land surface reflectance,land surface temperature,large-area mosaic image,and standard image map products,are described.The resulting Landsat RTU products are hosted on the RSGS earth observation data sharing web site for free download(http://ids.ceode.ac.cn/rtu/).These new products will provide consistent,standardized,multi-decadal image data for robust land cover change detection and monitoring across the Earth sciences.In the coming years,CASEarth DataBank system will be constructed,which is an intelligent data service platform for providing not only the RTU products from multi-source satellite data,but also big earth data analysis methods. | Guojin He Zhaoming Zhang Weili Jiao Tengfei Long Yan Peng Guizhou Wang Ranyu Yin Wei Wang Xiaomei Zhang Huichan Liu Bo Cheng Bo Xiang | 2018 | Big Earth Data2018,2,1: | 8 |
| 18 | Ontology-based Access Control for FAIR Data显示文摘This paper focuses on fine-grained,secure access to FAIR data,for which we propose ontology-based data access policies.These policies take into account both the FAIR aspects of the data relevant to access(such as provenance and licence),expressed as metadata,and additional metadata describing users.With this tripartite approach(data,associated metadata expressing FAIR information,and additional metadata about users),secure and controlled access to object data can be obtained.This yields a security dimension to the“A”(accessible)in FAIR,which is clearly needed in domains like security and intelligence.These domains need data to be shared under tight controls,with widely varying individual access rights.In this paper,we propose an approach called Ontology-Based Access Control(OBAC),which utilizes concepts and relations from a data set's domain ontology.We argue that ontology-based access policies contribute to data reusability and can be reconciled with privacy-aware data access policies.We illustrate our OBAC approach through a proof-of-concept and propose that OBAC to be adopted as a best practice for access management of FAIR data. | Christopher Brewster Barry Nouwt Stephan Raaijmakers Jack Verhoosel | 2020 | Data Intelligence2020,2,1: | 8 |
| 19 | Smart Data for Digital Humanities显示文摘The emergence of 'Big Data' has been a dramatic development in recent years.Alongside it,a lesser-known but equally important set of concepts and practices has also come into being—'Smart Data.' This paper shares the author's understanding of what,why,how,who,where,and which data in relation to Smart Data and digital humanities.It concludes that,challenges and opportunities co-exist,but it is certain that Smart Data,the ability to achieve big insights from trusted,contextualized,relevant,cognitive,predictive,and consumable data at any scale,will continue to have extraordinary value in digital humanities. | Marcia Lei Zeng | 2017 | Journal of Data and Information Science2017,2,1: | 8 |
| 20 | FAIR Science for Social Machines: Let’s Share Metadata Knowlets in the Internet of FAIR Data and Services显示文摘In a world awash with fragmented data and tools,the notion of Open Science has been gaining a lot of momentum,but simultaneously,it caused a great deal of anxiety.Some of the anxiety may be related to crumbling kingdoms,but there are also very legitimate concerns,especially about the relative role of machines and algorithms as compared to humans and the combination of both(i.e.,social machines).There are also grave concerns about the connotations of the term“open”,but also regarding the unwanted side effects as well as the scalability of the approaches advocated by early adopters of new methodological developments.Many of these concerns are associated with mind-machine interaction and the critical role that computers are now playing in our day to day scientific practice.Here we address a number of these concerns and provide some possible solutions.FAIR(machine-actionable)data and services are obviously at the core of Open Science(or rather FAIR science).The scalable and transparent routing of data,tools and compute(to run the tools on)is a key central feature of the envisioned Internet of FAIR Data and Services(IFDS).Both the European Commission in its Declaration on the European Open Science Cloud,the G7,and the USA data commons have identified the need to ensure a solid and sustainable infrastructure for Open Science.Here we first define the term FAIR science as opposed to Open Science.In FAIR science,data and the associated tools are all Findable,Accessible under well defined conditions,Interoperable and Reusable,but not necessarily“open”;without restrictions and certainly not always“gratis”.The ambiguous term“open”has already caused considerable confusion and also opt-out reactions from researchers and other data-intensive professionals who cannot make their data open for very good reasons,such as patient privacy or national security.Although Open Science is a definition for a way of working rather than explicitly requesting for all data to be available in full Open Access, the connotation of openness of the data involved in Open Science is very strong. In FAIR science, data and the associated services to run all processes in the data stewardship cycle from design of experiment to capture to curation, processing, linking and analytics all have minimally FAIR metadata, which specify the conditions under which the actual underlying research objects are reusable, first for machines and then also for humans. This effectively means that-properly conducted- Open Science is part of FAIR science. However, FAIR science can also be done with partly closed, sensitive and proprietary data. As has been emphasized before, FAIR is not identical to “open”. In FAIR/Open Science, data should be as open as possible and as closed as necessary. Where data are generated using public funding, the default will usually be that for the FAIR data resulting from the study the accessibility will be as high as possible, and that more restrictive access and licensing policies on these data will have to be explicitly justified and described. In all cases, however, even if the reuse is restricted, data and related services should be findable for their major uses, machines, which will make them also much better findable for human users. With a tendency to make good data stewardship the norm, a very significant new market for distributed data analytics and learning is opening and a plethora of tools and reusable data objects are being developed and released. These all need FAIR metadata to be routed to each other and to be effective. | Barend Mons | 2019 | Data Intelligence2019,1,1: | 8 |