Skip to main content

Home/ Data Working Group/ Group items tagged data-preservation

Rss Feed Group items tagged

Amy West

Interagency Data Stewardship/Citations/provider guidelines - Federation of Earth Scienc... - 0 views

    • Amy West
       
      Little confused by what's meant by "data sets should be cited like books" since they go on to provide really good reasons why data aren't like books, e.g. need subsetting information, access date for dynamic databases.
  • The guidelines build from the IPY Guidelines and are compatible with the DataCite Metadata Scheme for the Publication and Citation of Research Data, Version 2.2, July 2011.
  • In some cases, the data set authors may have also published a paper describing the data in great detail. These sort of data papers should be encouraged, and both the paper and the data set should be cited when the data are used.
  • ...27 more annotations...
  • Ongoing updates to a time series do change the content of the data set, but they do not typically constitute a new version or edition of a data set. New versions typically reflect changes in sampling protocols, algorithms, quality control processes, etc. Both a new version and an update may be reflected in the release date.
  • Locator, Identifier, or Distribution Medium
  • Then it is necessary to include a persistant reference to the location of the data.
  • This may be the most challenging aspect of data citation. It is necessary to enable "micro-citation" or the ability to refer to the specific data used--the exact files, granules, records, etc.
  • Data stewards should suggest how to reference subsets of their data. With Earth science data, subsets can often be identified by referring to a temporal and spatial range.
  • A particular data set may be part of a compilation, in which case it is appropriate to cite the data set somewhat like a chapter in an edited volume.
  • Increasingly, publishers are allowing data supplements to be published along with peer-reviewed research papers. When using the data supplement one need only cite the parent reference. F
  • Confusingly, a Digital Object Identifier is a locator. It is a Handle based scheme whereby the steward of the digital object registers a location (typically a URL) for the object. There is no guarantee that the object at the registered location will remain unchanged. Consider a continually updated data time series, for example.
  • While it is desirable to uniquely identify the cited object, it has proven extremely challenging to identify whether two data sets or data files are scientifically identical.
  • At this point, we must rely on location information combined with other information such as author, title, and version to uniquely identify data used in a study.
  • The key to making registered locators, such as DOIs, ARKS, or Handles, work unambiguously to identify and locate data sets is through careful tracking and documentation of versions.
  • how to handle different data set versions relative to an assigned locator.
  • Track major_version.minor_version.[archive_version].
  • Typically, something that affects the whole data set like a reprocessing would be considered a major version.
  • Assign unique locators to major versions.
  • Old locators for retired versions should be maintained and point to some appropriate web site that explains what happened to the old data if they were not archived.
  • A new major version leads to the creation of a new collection-level metadata record that is distributed to appropriate registries. The older metadata record should remain with a pointer to the new version and with explanation of the status of the older version data.
  • Major and minor version should be listed in the recommended citation.
  • inor versions should be explained in documentation
  • Ongoing additions to an existing time series need not constitute a new version. This is one reason for capturing the date accessed when citing the data.
  • we believe it is currently impossible to fully satisfy the requirement of scientific reproducibility in all situations
  • To aid scientific reproducibility through direct, unambiguous reference to the precise data used in a particular study. (This is the paramount purpose and also the hardest to achieve). To provide fair credit for data creators or authors, data stewards, and other critical people in the data production and curation process. To ensure scientific transparency and reasonable accountability for authors and stewards. To aid in tracking the impact of data set and the associated data center through reference in scientific literature. To help data authors verify how their data are being used. To help future data users identify how others have used the data.
  • The ESIP Preservation and Stewardship cluster has examined these and other current approaches and has found that they are generally compatible and useful, but they do not entirely meet all the purposes of Earth science data citation.
  • In general, data sets should be cited like books.
  • hey need to use the style dictated by their publishers, but by providing an example, data stewards can give users all the important elements that should be included in their citations of data sets
  • Access Date and Time--because data can be dynamic and changeable in ways that are not always reflected in release dates and versions, it is important to indicate when on-line data were accessed.
  • Additionally, it is important to provide a scheme for users to indicate the precise subset of data that were used. This could be the temporal and spatial range of the data, the types of files used, a specific query id, or other ways of describing how the data were subsetted.
Lisa Johnston

Chronopolis -- Digital Preservation Program -- Long-Term Mass-Scale Federated Digital P... - 0 views

  •  
    The Chronopolis Digital Preservation Demonstration Project, one of the Library of Congress' latest efforts to collect and preserve at-risk digital information, has been officially launched as a multi-member partnership to meet the archival needs of a wide range of cultural and social domains. Chronopolis is a digital preservation data grid framework being developed by the San Diego Supercomputer Center (SDSC) at UC San Diego , the UC San Diego Libraries (UCSDL) , and their partners at the National Center for Atmospheric Research (NCAR) in Colorado and the University of Maryland's Institute for Advanced Computer Studies (UMIACS) . A key goal of the Chronopolis project is to provide cross-domain collection sharing for long-term preservation. Using existing high-speed educational and research networks and mass-scale storage infrastructure investments, the partnership is designed to leverage the data storage capabilities at SDSC, NCAR, and UMIACS to provide a preservation data grid that emphasizes heterogeneous and highly redundant data storage systems.
Amy West

Open access to research data a lot tougher than you think - 2 views

  • It means that researchers need to deal with the formatting and deposition of data, an annoying step when they would rather be focusing on their next project. Given the time lag, it's also difficult to associate the correct metadata with the material that's being a
  • According to the commentary, scientists view data deposition as a burden due to the extra work it involves. Research data is usually not in the correct format for submission to repositories when the project is completed, and so the scientist must take the time to convert it.
  • The authors here propose a new approach to data management, where each research institution should employ data managers to work with scientists and administer local, structured data storage. Local storage and support is the preference of most scientists, who would rather not hand off control of their data to remote strangers.
Amy West

2011AGUworkshop - Federation of Earth Science Information Partners - 1 views

  •  
    All the presentations are good, but I found the Data formats, Creating documentation & metadata, working w/an archive & preservation strategies particularly good. Solid examples of formats, metadata, and real-life preservation. Plus, as mgs of UDC/AgEcon, hopefully more archives over time, I think we should look hard at what they tell researchers to look for in an archive.
Lisa Johnston

Digital Curation Centre: DCC SCARP Project - 0 views

  •  
    18 January 2010 | Key perspectives | Type: report The Digital Curation Centre is pleased to announce the report "Data Dimensions: Disciplinary Differences in Research Data Sharing, Reuse and Long term Viability" by Key Perspectives, as one of the final outputs of the DCC SCARP project. The project investigated attitudes and approaches to data deposit, sharing and reuse, curation and preservation, over a range of research fields in differing disciplines. The synthesis report (which drew on the SCARP case studies plus a number of others, identified in the Appendix), identifies factors that help understand how curation practices in research groups differ in disciplinary terms. This provides a backdrop to different digital curation approaches.
Amy West

Data Preservation - Home - 1 views

  •  
    USGS is attempting to corral / manage geological & geophysical preservation efforts.
Lisa Johnston

Sustainable Digital Preservation and Access - 0 views

  •  
    While storage and technological issues have been at the forefront of the discussion on digital information, relatively little focus has been on the economic aspect of preserving vast amounts of digital data fundamental to the modern world.
Lisa Johnston

Data Preservation and Policy - 11 views

Policy issues at the national level are of particular interest to those involved with e-science initiatives. Several organization have emerged at the forefront of this arena and it will be particul...

policy national mandate grant

started by Lisa Johnston on 05 Nov 08 no follow-up yet
Lisa Johnston

San Diego Supercomputer Center director offers tips on data preservation in the informa... - 0 views

  •  
    Communications of the ACM,
Lisa Johnston

Geospatial Data Preservation - View Resources Tools & Software - 2 views

  •  
    Really nice resource of research articles, tools, and policies for geo data and beyond.
Lisa Johnston

UNM Today: University Libraries Hosts Lecture on Preserving and Using Large Data Sets - 0 views

  •  
    Univ New Mexico press release on their NSF funded DataNet project
1 - 14 of 14
Showing 20 items per page