Try a new search

Format these results:

Searched for:

in-biosketch:yes

person:sbj2002

Total Results:

132


Assessing data relevance for automated generation of a clinical summary

Van Vleck, Tielman T; Stein, Daniel M; Stetson, Peter D; Johnson, Stephen B
Clinicians perform many tasks in their daily work requiring summarization of clinical data. However, as technology makes more data available, the challenges of data overload become ever more significant. As interoperable data exchange between hospitals becomes more common, there is an increased need for tools to summarize information. Our goal is to develop automated tools to aid clinical data summarization. Structured interviews were conducted on physicians to identify information from an electronic health record they considered relevant to explaining the patients medical history. Desirable data types were systematically evaluated using qualitative and quantitative analysis to assess data categories and patterns of data use. We report here on the implications of these results for the design of automated tools for summarization of patient history.
PMCID:2655814
PMID: 18693939
ISSN: 1942-597x
CID: 3586242

An unsupervised machine learning approach to segmentation of clinician-entered free text

Wrenn, Jesse O; Stetson, Peter D; Johnson, Stephen B
Natural language processing, an important tool in biomedicine, fails without successful segmentation of words and sentences. Tokenization is a form of segmentation that identifies boundaries separating semantic units, for example words, dates, numbers and symbols, within a text. We sought to construct a highly generalizeable tokenization algorithm with no prior knowledge of characters or their function, based solely on the inherent statistical properties of token and sentence boundaries. Tokenizing clinician-entered free text, we achieved precision and recall of 92% and 93%, respectively compared to a whitespace token boundary detection algorithm. We classified over 80% of punctuation characters correctly, based on manual disambiguation with high inter-rater agreement (kappa=0.916). Our algorithm effectively discovered properties of whitespace and punctuation in the corpus without prior knowledge of either. Given the dynamic nature of biomedical language, and the variety of distinct sublanguages used, the effectiveness and generalizability of our novel tokenization algorithm make it a valuable tool.
PMCID:2655800
PMID: 18693949
ISSN: 1942-597x
CID: 3586252

Feasibility study of speech recognition for gathering information needs

Natarajan, Karthik; Duffy, Robert F; Johnson, Stephen B; Mendonça, Eneida A
Automated speech recognition (ASR) is used in many areas of medicine today. However, not many studies have evaluated the usefulness of ASR applications for capturing clinician information needs in noisy environments. We evaluated 72 ASR transcribed clinician-generated questions and assessed them for semantic and syntactic errors. The results showed that basic user training is not sufficient in order to capture the semantics of recordings.
PMID: 18694157
ISSN: 1942-597x
CID: 3586262

Graph theoretic modeling of large-scale semantic networks

Bales, Michael E; Johnson, Stephen B
During the past several years, social network analysis methods have been used to model many complex real-world phenomena, including social networks, transportation networks, and the Internet. Graph theoretic methods, based on an elegant representation of entities and relationships, have been used in computational biology to study biological networks; however they have not yet been adopted widely by the greater informatics community. The graphs produced are generally large, sparse, and complex, and share common global topological properties. In this review of research (1998-2005) on large-scale semantic networks, we used a tailored search strategy to identify articles involving both a graph theoretic perspective and semantic information. Thirty-one relevant articles were retrieved. The majority (28, 90.3%) involved an investigation of a real-world network. These included corpora, thesauri, dictionaries, large computer programs, biological neuronal networks, word association networks, and files on the Internet. Twenty-two of the 28 (78.6%) involved a graph comprised of words or phrases. Fifteen of the 28 (53.6%) mentioned evidence of small-world characteristics in the network investigated. Eleven (39.3%) reported a scale-free topology, which tends to have a similar appearance when examined at varying scales. The results of this review indicate that networks generated from natural language have topological properties common to other natural phenomena. It has not yet been determined whether artificial human-curated terminology systems in biomedicine share these properties. Large network analysis methods have potential application in a variety of areas of informatics, such as in development of controlled vocabularies and for characterizing a given domain.
PMID: 16442849
ISSN: 1532-0480
CID: 3586072

Is the Health Level 7/LOINC document ontology adequate for representing nursing documents?

Hyun, Sookyung; Ventura, Rosemary; Johnson, Stephen B; Bakken, Suzanne
The use of nursing documents from different electronic health record (EHR) systems is challenging due to inconsistency in document naming across systems and institutions. Mapping each local document name to standard document ontology may enable health care professionals to navigate and retrieve documents efficiently for multiple purposes such as quality assurance, outcomes research or public health reporting. The purpose of this study was to evaluate the sufficiency of the Health Level 7 (HL7)/Logical Observation Identifiers, Names, and Codes (LOINC) document ontology for representing nursing document names. We collected 94 nursing document types from the Eclipsys Clinical Information System (CIS) and the Columbia Medical Entities Dictionary (MED) and mapped them to the components of the HL7/LOINC document ontology. Seventy-five (79.8%) nursing document names were completely represented and 19 (20.2%) document names were partially represented. In order for the HL7/LOINC document ontology to be of more use in implementing EHRs that support nursing documentation, Subject Matter Domain and Type of Service axes require extension and clarification.
PMID: 17102314
ISSN: 0926-9630
CID: 3586122

Markup of temporal information in electronic health records

Hyun, Sookyung; Bakken, Suzanne; Johnson, Stephen B
Temporal information plays a critical role in the understanding of clinical narrative (i.e., free text). We developed a representation for marking up temporal information in a narrative, consisting of five elements: 1) reference point, 2) direction, 3) number, 4) time unit, and 5) pattern. We identified 254 temporal expressions from 50 discharge summaries and represented them using our scheme. The overall inter-rater reliability among raters applying the representation model was 75 percent agreement. The model can contribute to temporal reasoning in computer systems for decision support, data mining, and process and outcomes analyses by providing structured temporal information.
PMID: 17102457
ISSN: 0926-9630
CID: 3586132

Reengineering clinical research with informatics

Chung, Thomas K; Kukafka, Rita; Johnson, Stephen B
The future success of the translational research spectrum depends on the clinical research enterprise's ability to break through the barriers that constrain its productivity. As more basic science discoveries emerge, our ability to effectively translate this knowledge into improved patient care rests squarely on the manner in which we answer clinical questions. Informatics--the science of effective information use--is poised to help advance the conduct of science. However, incorporating informatics into the enterprise comes with its own set of challenges. To harness the benefits of improved information use, it is important to first establish how information flows within research. A thoughtful implementation of informatics--one that factors in social and organizational nuances--will undoubtedly lead to a more efficient and effective clinical research enterprise.
PMID: 17134616
ISSN: 1081-5589
CID: 3586142

Modeling clinical trials workflow in community practice settings

Khan, Sharib A; Payne, Philip R O; Johnson, Stephen B; Bigger, J Thomas; Kukafka, Rita
Clinical research is vital to the translation of biomedical knowledge into standard clinical practice. Efforts are underway under the NIH Roadmap initiative to re-engineer the national research enterprise to sustain the rapid pace of innovation in the biomedical domain. As part of these efforts, we have embarked on an empirical evaluation of clinical research workflow in community practice settings. The reasons for this focus are three-fold. First, there is an increasing tendency by trial sponsors to conduct clinical trials in community, rather than academic, settings. Second, understanding workflow is critical to developing re-engineering strategies. Third, workflow associated with the conduct of clinical research in community practices have received virtually no attention in the scientific literature. In this paper, we describe a pilot study using time-motion observations, to determine the workflow of clinical research coordinators, the tools they use to conduct the constituent activities of those workflows, and their ultimate outcomes. The preliminary findings provide insights and understanding of clinical research workflow in community practice settings - knowledge that may significantly impact the way in which information technology based re-engineering can be deployed in such an environment.
PMCID:1839500
PMID: 17238375
ISSN: 1942-597x
CID: 3586152

Heuristic evaluation of eNote: an electronic notes system

Bright, Tiffani J; Bakken, Suzanne; Johnson, Stephen B
eNote is an electronic health record (EHR) system based on semi-structured narrative documents. A heuristic evaluation was conducted with a sample of five usability experts. eNote performed highly in: 1)consistency with standards and 2)recognition rather than recall. eNote needs improvement in: 1)help and documentation, 2)aesthetic and minimalist design, 3)error prevention, 4)helping users recognize, diagnosis, and recover from errors, and 5)flexibility and efficiency of use. The heuristic evaluation was an efficient method of evaluating our interface.
PMCID:1839434
PMID: 17238484
ISSN: 1942-597x
CID: 3586162

Modeling electronic discharge summaries as a simple temporal constraint satisfaction problem

Hripcsak, George; Zhou, Li; Parsons, Simon; Das, Amar K; Johnson, Stephen B
OBJECTIVE:To model the temporal information contained in medical narrative reports as a simple temporal constraint satisfaction problem. DESIGN/METHODS:A constraint satisfaction problem is defined by time points and constraints (inequalities between points). A time interval comprises a pair of points and a constraint. Five complete electronic discharge summaries and paragraphs from 226 other discharge summaries were studied. Medical events were represented as intervals, and assertions about events were represented as constraints. Through a consensus process, a set of encoding procedures and a list of issues related to encoding were generated. MEASUREMENTS/METHODS:Instances of temporal disjunction and contradiction and distribution of temporal constraints were used. RESULTS:An average of 95 medical events (range, 46-151) and 234 temporal assertions (range, 118-388) were identified per complete discharge summary. Nondefinitional assertions were explicit (36%) or implicit (64%) and absolute (17%), qualitative (72%), or metric (11%). Implicit assertions were based on domain knowledge and assumptions, e.g., the section of the report determined the ordering of events. Issues included linking events, intermittence, periodicity, granularity, vagueness, ambiguity, uncertainty, and plans. ions such as intermittence were not represented explicitly. The temporal network was sparse: Only 0.80% (range, 0.42%-1.38%) of possible constraints were instantiated. No instances of discontinuous temporal disjunction were found in the complete summaries or the 226 paragraphs. One instance of temporal contradiction was found (intrareport rate of 0.2 with a 95% confidence interval of 0.005-1.114). CONCLUSION/CONCLUSIONS:A simple temporal constraint satisfaction problem appears sufficient to represent most temporal assertions in discharge summaries and may be useful for encoding electronic medical records.
PMCID:543827
PMID: 15492038
ISSN: 1067-5027
CID: 3585972