This illustrated that with some degree of experience, the users, particularly the more knowledgeable users, found the use of ontology to be more useful
This illustrated that with some degree of experience, the users, particularly the more knowledgeable users, found the use of ontology to be more useful. A number of users pointed out that it will be beneficial if they could (a) have the option to select the sources they wanted to search over, and (b) specify the type of results they wanted (e.g., results with image content only). A number of users showed cases where the NIF Web retrieves some data pages that are not in the domain of neuroscience. The users almost unanimously stated that they wanted the results of the queries to be organized by ontological terms instead of (or in addition to) by resource type. achieve this functionality. A central element in the system is an ontology called NIFSTD (for NIF Standard) constructed by amalgamating a number of known and newly developed ontologies. NIFSTD is used by our ontology management module, called OntoQuest to perform ontology-based search over data sources. The NIF architecture PF-06855800 currently provides three different mechanisms for searching heterogeneous data sources including relational databases, web sites, XML documents and full text of publications. Version 1.0 of the PF-06855800 NIF system is currently in beta test and may be accessed throughhttp://nif.nih.gov. Keywords:ontology, data federation, neuroscience resource == Introduction == Today, there are thousands of neuroscience information resources created by a wide range of information providers including research groups, funding agencies, vendor groups and public data initiatives that publish information in one form or another. A neuroscience information resource is any electronically accessible site that provides information of interest to neuroscience. A neuroscience information resource can be a digital library of publications like PubMed; it can be the web site of a neuroscience research group that publishes its research detail as web pages; it can be a tissue bank that allows a potential neuroscientist customer to navigate through its samples; it can be a database that houses experimental research results, PF-06855800 and allows users to query it; it can even be a software tool that enables a user to perform a computation online. Unfortunately, despite the growing body of information resources, the problem of finding just the right information from one or more of these has not become easier, and it may be very hard for a general neuroscientist to locate a relevant information resource if she does not know about its existence. Let us consider the following example to illustrate the problem. Assume that a neuroscientist is looking for resources that might provide cDNA for mouse PF-06855800 models. Typically, she would use the Google search engine with the keyword combination like mouse, model, cDNA.Figure 1(a)shows the first result page returned by Google. Although the results are indeed about mouse models and cDNA, they are mostly URLs of Google-indexed publications and general web sites. == Figure 1. == Figure 1(a). The result of the query (mouse model cDNA) against Google. The top results are very general, and mostly from papers that are indexed by Google. Figure 1(b). The same query as in Figure 1(a) now executed against NIF. In contrast with Google, the selective web crawling coverage of NIF enables it to return results that are more closely related to Neuroscience. A more focused search that might actually help the neuroscientist better is shown inFigure 1(b). This result is mostly about resources like Open Biosystems that can be used as a resource for cDNA libraries for mouse models of human diseases. The resource finding problem gets compounded if the neuroscientist wants to search for the information not only from the web, but overanykind of information resource mentioned in the previous paragraph, because there are no search tools that provide adequate functionality to satisfy the information needs of our neuroscientist. == The Problem == There are a number of underlying factors behind the resource finding problem. These factors are not specific to the domain of neuroscience, but come into play whenever a discipline-specific information seeker tries to locate information resources that have been created for very different goals, have heterogeneous content, provide heterogeneous access mechanisms, and have not been put into a common information framework. In the following, we list the contributing factors that need to be overcome to allow an information seeker find meaningful results quickly over these heterogeneous data sources. Although a domain user searches for resources using keywords, the intent of the search isconceptual. Thus, although the search termastrocytoma, there is an implicit expectation that a resource aboutastrocytic gliomaorglial malignancywill be part of the result. Most search engines do not provide asemantic search facility, which includes not only search by synonyms but by terms that are notionally related to the search terms. As another example, the user searching forhippocampal formationwould possibly also be interested in the cell types found Rabbit Polyclonal to ARSA there because they are semantically related. The web is an important class of information resource, but.