Join the 2020 ESIP Winter Meeting Highlights Webinar on Feb. 5th at 3 pm ET for a fast-paced overview of what took place at the meeting. More info here.

Tuesday, January 7

11:00am EST

Creating a Data at Risk Commons at DataAtRisk.org
Several professional organizations have become increasingly concerned about the loss of reusable data from primary sources such as individual researchers, projects, and agencies. DataAtRisk.org aims to connect people with data in need, to data expertise, and is a response to the clear need for a community building application. This “Data at Risk” commons will allow individuals to submit and request help with threatened datasets and connect these datasets to experts who can provide resources and skills to help rescue data through a secure, professional mechanism to facilitate self-identification and discovery.

This session will provide an overview of the current status of the DataAtRisk.org project, and aims to expand the network of individuals involved in the development and implementation of DataAtRisk.org

How to Prepare for this Session: Please check out https://dataatrisk.org/ for some background on the activities.

Presentations: http://bit.ly/303gig7, https://doi.org/10.6084/m9.figshare.11536317.v1
Link to use case / user scenario: https://tinyurl.com/yh4rnk7b

View Recording: https://youtu.be/96NMQwx_EtI

  • Perfection is the enemy of getting stuff done
  • Something is better than nothing
  • Triage will be necessary at several places in the process

Director, Energy Investigations, Geological Survey of Alabama
Long tail data, data preservation, connecting physical samples to digital information, geoscience policy, science communication

Tuesday January 7, 2020 11:00am - 12:30pm EST
Linden Oak
  Working Session

4:00pm EST

Experiences Migrating Mission Scale Data in the Cloud
We will describe our project to upload a 2.4 PB dataset encapsulated into ~80K fused files from the 5 instruments on the Terra satellite into NASA AWS S3.
We will share the bottlenecks points and lessons learned during this process and expect to share experiences with similar projects in order to understand the best practices and collect guidelines for future projects that are adopting cloud solutions for their data needs.

We'll discuss data volumes, data integrity strategies for migration, S3 bucket organization, metadata curation, transfer rates, transfer pipelines, etc. We will also discuss and share data access patterns, costs, and architectures and how we can construct guidelines for access to these datasets efficiently.

We encourage the discussion among different projects that faced similar processes or are looking to migrate their datasets into the cloud.



View Recording: https://youtu.be/1xVJghJI4Gg

  • Project required/used a combination of NSF, NASA and AWS resources. Some interesting discussion around AWS or other cloud services as a stand in or follow on to limited term NSF assets
  • Some interesting discussion of tailoring to appropriate end users- wide range of potential users and thus requirements for the dataset. This includes access guidelines, user capabilities etc.
  • Project aimed to make a paradigm shift from understanding/observing physical processes to a full climate observing objective

Research Programmer, National Center for Supercomputing Applications Connect Message

Tuesday January 7, 2020 4:00pm - 5:30pm EST
White Flint
  Breakout
Wednesday, January 8

2:00pm EST

Advancing Data Integration approaches of the structured data web
Political, economic, social or scientific decision making is often based on integrated data from multiple sources across potentially many disciplines. To be useful, data need to be easy to discover and integrate.
This session will feature presentations highlighting recent breakthroughs and lessons learned from experimentation and implementation of open knowledge graph, linked data concepts and Discrete Global Grid Systems. Practicality and adoptability will be the emphasis - focusing on incremental opportunities that enable transformational capabilities using existing technologies. Best practices from the W3C Spatial Data on the Web Working Group, OGC Environmental Linked Features Interoperability Experiment, ESIP Science on Schema.org; implementation examples from Geoscience Australia, Ocean Leadership Consortium, USGS and other organisations will featured across the entire session.
This session will highlight how existing technologies and best practices can be combined to address important and common use cases that have been difficult if not impossible until recent developments. A follow up session will be used to seed future collaborative development through co-development, github issue creation, and open documentation generation.

How to Prepare for this Session: Review: https://opengeospatial.github.io/ELFIE/, https://github.com/ESIPFed/science-on-schema.org, https://www.w3.org/TR/sdw-bp/, and http://locationindex.org/.

Notes, links, and attendee contact info here.

View Recording: https://youtu.be/-raMt2Y1CdM

Session Agenda:
1.  2.00- 2.10,  Sylvain Grellet, Abdelfettah Feliachi, BRGM, France
'Linked data' the glue within interoperable information systems
“Our Environmental Information Systems are exposing environmental features, their monitoring systems and the observation they generate in an interoperable way (technical and semantic) for years. In Europe, there is even a legal obligation to such practices via the INSPIRE directive. However, the practice inducing data providers to set up services in a "Discovery > View > Download data" pattern hides data behind the services. This hinders data discovery and reuse. Linked Data on the Web Best Practices put this stack upside down and data is now back in the first line. This completely revamp the design and capacities of our Information Systems. We'll highlight the new data frontiers opened by such practices taking examples on the French National Groundwater Information Network”
View Slides: https://doi.org/10.6084/m9.figshare.11550570.v1

2.  2.10 - 2.20,  Adam Leadbetter, Rob Thomas, Marine Institute, Ireland
Using RDF Data Cubes for data visualization: an Irish pilot study for publishing environmental data to the semantic web
The Irish Wave and Weather Buoy Networks return metocean data at 5-60 minute intervals from 9 locations in the seas around Ireland. Outside of the Earth Sciences an example use case for these data is in supporting Blue Economy development and growth (e.g. renewable energy device development). The Marine Institute, as the operator of the buoy platforms, in partnership with the EU H2020 funded Open Government Intelligence project has published daily summary data from these buoys using the RDF DataCube model[1]. These daily statistics are available as Linked Data via a SPARQL endpoint making these data semantically interoperable and machine readable. This API underpins a pilot dashboard for data exploration and visualization. The dashboard presents the user with the ability to explore the data and derive plots for the historic summary data, while interactively subsetting from the full resolution data behind the statistics. Publishing environmental data with these technologies makes accessing environmental data available to developers outside those with Earth Science involvement and effectively lowers the entry bar for usage to those familiar with Linked Data technologies.
View Slides: https://doi.org/10.6084/m9.figshare.11550570.v1

3. 2.20 - 2.30,  Boyan Brodaric, Eric Boisvert, Geological Survey of Canada, Canada; David Blodgett, USGS, USA
Toward a Linked Water Data Infrastructure for North America
We will describe progress on a pilot project using Linked Data approaches to connect a wide variety of water-related information within Canada and the US, as well as across the shared border
View Slides: https://doi.org/10.6084/m9.figshare.11541984.v1

4.  2.30 - 2.40,  Dalia Varanka, E. Lynn Usery, USGS, USA
The Map as Knowledge Base; Integrating Linked Open Topographic Data from The National Map of the U.S. Geological Survey
This presentation describes the objectives, models, and approaches for a prototype system for cross-thematic topographic data integration based on semantic technology. The system framework offers a new perspectives on conceptual, logical, and physical system integration in contrast to widely used geographic information systems (GIS).
View Slides: https://doi.org/10.6084/m9.figshare.11541615.v1

5.  2.40 – 2.50,  Alistair Ritchie, Landcare, New Zealand
ELFIE at Landcare Research, New Zealand
Landcare Research, a New Zealand Government research institute, creates, manages and publishes a large set of observational and modelling data describing New Zealand’s land, soil, terrestrial biodiversity and invasive species. We are planning to use the findings of the ELFIE initiatives to guide the preparation of a default view of the data to help discovery (by Google), use (by web developers) and integration (into the large environmental data commons managed by other agencies). This integration will not only link data about the environment together, but will also expose more advanced data services. Initial work is focused on soil observation data, and the related scientific vocabularies, but we anticipate near universal application across our data holdings.
View Slides: https://doi.org/10.6084/m9.figshare.11550369.v1

6.  2.50 - 3.00,  Irina Bastrakova, Geoscience Australia, Australia
Location Index Project (Loc-I) – integration of data on people, business & the environment
Location Index (Loc-I) is a framework that provides a consistent way to seamlessly integrate data on people, business, and the environment.
Location Index aims to extend the characteristics of the foundation spatial data of taking geospatial data (multiple geographies) which is essential to support public safety and wellbeing, or critical for a national or government decision making that contributes significantly to economic, social and environmental sustainability and linking it with observational data. Through providing the infrastructure to suppo

Wednesday January 8, 2020 2:00pm - 3:30pm EST
White Flint

4:00pm EST

Structured data web and coverages integration working session
This working session will follow on the "Advancing Data Integration approaches of the structured data web” session and the Coverage Analytics sprint as an opportunity for those interested in building linked data information products that integrate spatial features, coverage data, and more. As such, inspiration will be drawn from projects like science on schema.org, the Environmental Linked Features Interoperability Experiment, the Australian Location Index, and those that session attendees take part in. Participants will self organize into use-case or technology focused groups to discuss and synthesize the outcomes of the sprint and structured data web session. Session outcomes could take a number of forms: linked data and web page mock ups, ideas and issues for OGC, W3C, or ESIP groups to consider, example data or use cases for relevant software development projects to consider, or work plans and proposals for suture ESIP work. The session format is expected to be fluid with an ideation and group formation exercise followed by structured discussion to explore a set of ideas then narrow on a focused valuable outcome. Participants will be encouraged to work together prior to the meeting to design and plan the session structure. Outcomes of the session will be reported at an Information Technology and Interoperability webinar in early 2020. How to Prepare for this Session: Attend the coverage sprint and the "Advancing Data Integration approaches of the structured data web" session.

Shared document for session here.

Full Notes: https://doi.org/10.6084/m9.figshare.11559087.v1


View Recording: https://youtu.be/u2x3I0cr46A

  • Takeaways
    Breakout session information interoperability committee and webinar series. See notes: https://docs.google.com/document/d/1LpcTMwP0mAD4G4Gb8mStI5uSDV61_qWPUkQ9nI1x1cI/edit?usp=sharing
  • Foster cross-project consistency via breakouts. Such as dealing with science on schema.org issue of Links to “in-band” linked (meta)data and “out of band” linked data. Content negotiation and in-band and out of band links Use blank nodes with link properties for rdf elements that are URI for out of band content. Identify in band links with sdo @id, out of band links with sdo:URL
  • Incorporating Spatial Coverages in Knowledge Graphs; Next Steps? Need to explore more on tessellations as an intermediate index. Will carry forward some of these ideas at the EDR SWG Will represent some of these ideas to the OGC-API Coverages SWG Will mention these ideas to the UFOKN Role of ‘spatial’ knowledge graphs Will spatial data analysis and transformation tools grow to adopt/support RDF as an underlying data structure for spatial information or will RDF continue to be a ‘view’ of existing (legacy) spatial data in GI systems?

Wednesday January 8, 2020 4:00pm - 5:30pm EST
White Flint
Thursday, January 9

10:15am EST

Working Group for the Data Stewardship Committee
This session is a working group for the 2020-2021 year for the Data Stewardship committee. We will discuss priorities for the next year, potential collaborative outputs, and review the work in progress from the last year. 

Notes Document: https://docs.google.com/document/d/1B_0K5jGnFgH72U3P2-oGr5vEqHOGU8CWU-IkZ6pjXbM/edit?ts=5e174588


View Recording: https://youtu.be/am-ZLfHgM4w

  • Wow, the members of the Committee really are active! Practically everyone has their own cluster or two!
  • Six activities proposed for the upcoming year have champions who will lead the effort to define the outputs of their selected activity.

Thursday January 9, 2020 10:15am - 11:45am EST
Forest Glen
  Business Meeting

12:00pm EST

License Up! What license works for you and your downstream repositories?
Many repositories are seeing an increase in the use and diversity of licenses and other intellectual property management (IPM) tools applied to externally-created data submissions and software developed by staff. However, adding a license to data files may have unexpected or unintended consequences in the downstream use or redistribution of those data. Who “owns” the intellectual property rights to data collected by university researchers using Federal and State (i.e., public) funding that must be deposited at a Federal repository? What license is appropriate for those data and what — exactly — does that license allow and disallow? What kind of license or other IPM instrument is appropriate for software written by a team of Federal and Cooperative Institute software engineers? Is there a significant difference between Creative Commons, GNU, and other ‘open source licenses’?

We have invited a panel of legal advisors from Federal and other organizations to discuss the implications of these questions for data stewards and the software teams that work collaboratively with those stewards. We may also discuss the latest information about Federal data licenses as it applies to the OPEN Government Data Act of 2019. How to Prepare for this Session: Consider what, if any, licenses, copyright, or other intellectual property rights management you apply or think applies to your work. Also consider Federal requirements such as the OPEN Government Data Act of 2019, Section 508 of the Rehabilitation Act of 1973.

Dr. Robert J. Hanisch is the Director of the Office of Data and Informatics, Material Measurement Laboratory, at the National Institute of Standards and Technology in Gaithersburg, Maryland. He is responsible for improving data management and analysis practices and helping to assure compliance with national directives on open data access. Prior to coming to NIST in 2014, Dr. Hanisch was a Senior Scientist at the Space Telescope Science Institute, Baltimore, Maryland, and was the Director of the US Virtual Astronomical Observatory. For more than twenty-five years Dr. Hanisch led efforts in the astronomy community to improve the accessibility and interoperability of data archives and catalogs.
Henry Wixon is Chief Counsel for the National Institute of Standards and Technology (NIST) of the U.S. Department of Commerce. His office provides programmatic legal guidance to NIST, as well as intellectual property counsel and representation to the Department of Commerce and other Department bureaus. In this role, it interacts with principal developers and users of research, including private and public laboratories, universities, corporations and governments. Responsibilities of Mr. Wixon’s office include review of NIST Cooperative Research and Development Agreements (CRADAs), licenses, Non-Disclosure Agreements (NDAs) and Material Transfer Agreements (MTAs), and the preparation and prosecution of the agency’s patent applications. As Chief Counsel, Mr. Wixon is active in standing Interagency Working Groups on Technology Transfer, on Bayh-Dole, and on Research Misconduct, as well as in the Federal Laboratory Consortium. He is a Certified Licensing Professional and a Past Chair of the Maryland Chapter of the Licensing Executives Society, USA and Canada (LES), and is a member of the Board of Visitors of the College of Computer, Mathematical and Natural Sciences of the University of Maryland, College Park.

See attached

View Recording: https://youtu.be/5Ng5FDW1LXk.



Thursday January 9, 2020 12:00pm - 1:30pm EST
Forest Glen
  Panel

12:00pm EST

Datacubes for Analysis-Ready Data: Standards & State of the Art
This workshop session will follow up on the OGC Coverage Analytics sprint, focusing specifically on advanced services for spatio-temporal datacubes. In the Earth sciences datacubes are accepted as an enabling paradigm for offering massive spatio-temporal Earth data analysis-ready, more generally: easing access, extraction, analysis, and fusion. Also, datacubes homogenizes APIs across dimensions, allowing unified wrangling of 1-D sensor data, 2-D imagery, 3-D x/y/t image timeseries and x/y/z geophysics voxel data, and 4-D x/y/z/t climate and weather data.
Based on the OGC datacube reference implementation we introduce datacube concepts, state of standardization, and real-life 2D, 3D, and 4D examples utilizing services from three continents. Ample time will be available for discussion, and Internet-connected participants will be able to replay and modify many of the examples shown. Further, key datacube activities worldwide, within and beyond Earth sciences, will be related to.
Session outcomes could take a number of forms: ideas and issues for OGC, ISO, or ESIP to consider; example use cases; challenges not yet addressed sufficiently, and entirely novel use cases; work and collaboration plans for future ESIP work. Outcomes of the session will be reported at the next OGC TC meeting's Big Data and Coverage sessions. How to Prepare for this Session: Introductory and advanced material is available from http://myogc.org/go/coveragesDWG


View Recording: https://youtu.be/82WG7soc5bk

  • Abstract coverage construct defines the base which can be filled up with a coverage implementation schema. Important as previously implementation wasn’t interoperable with different servers and clients. 
  • Have embedded the coordinate system retrieved from sensors reporting in real time into their xml schema to be able to integrate the sensor data into the broader system. Can deliver the data in addition to GML but JSON, and RDF which could be used to link into semantic web tech. 
  • Principle is send HTTP url-encoded query to server and get some results that are extracted from datacube, e.g., sources from many hyperspectral images.


Thursday January 9, 2020 12:00pm - 1:30pm EST
White Flint