The "SCUBA-2 All-Sky Survey" as a CANFAR project ------------------------------------------------ Abstract -------- The SCUBA-2 `All-Sky' Survey (SASSy) is an ambitious project to perform a high-resolution search of all of the sky observable by the JCMT on Mauna Kea, Hawaii at a wavelength of 850 um. The primary science aims are to probe star formation throughout the Universe, with specific goals to detect all nearby regions of star formation, investigate star formation at high Galactic latitudes, determine the numbers and distribution of infrared-dark clouds in the Galaxy as well as bright submillimetre star-forming external galaxies and to provide detailed foreground maps for the Planck Cosmic background satellite. Detailed case ------------- SASSy is an international collaboration of 48 astronomers from more than 20 institutions in Canada, the UK and Netherlands. The survey will be carried out using SCUBA-2 at the JCMT, with raw data being archived at the CADC. All subsequent data access will be through CADC. The large area to be covered by the survey requires easy comparison with existing data and astronomers will want to query existing source catalogues and compare SASSy images with those from previous observations made with SCUBA, CGPS, IRAS, CFHT and HST, all accessible through CADC. The size of the raw data, the required compute power for pipeline reduction, the need to develop new map-making tools, the distributed nature of the collaboration and the planned involvement of CADC make this an ideal project within CANFAR. SCUBA-2 generates of order 500 GB of raw data per night and the processing of the data into images is very computationally-intensive. For small projects, desktop processing is feasible but for a survey such as SASSy this is impossible and significant memory and disk-space resources are required. A large survey such as SASSy will benefit enormously from distributed processing of individual fields which are later combined into a single large mosaic. The size of the collaboration, and the scope of the project means that it is simply not feasible for processing to be carried out at any individual institution. A single, central processing and project management location is vital for the success of SASSy to ensure consistency and ease of management. The survey has multiple coordinators (with the Canadian one based at UBC) who will require a central place to track the progress of the survey. Other members of the collaboration will need to examine the latest data products for quality assessment and be able to approve data products or issue requests for re-processing. These requests (and any changes to the processing steps) must be recorded for all members of the collaboration to examine. SASSy will scan the sky as fast as the telescope can take data in order to cover the largest area possible. To save time, while facilitating a basic level of quality assurance, each patch of sky will normally be covered twice during the survey. Furthermore, given the time constraints the project will be scheduled to observe regions of the sky in appropriate weather conditions (e.g. regions at low elevation must be observed in better weather than region at high elevation). If an observation fails to satisfy the data quality criteria, it is imperative that a decision be made quickly to repeat that observation so that the survey can proceed as efficiently as possible. In addition, if a source is marginally detected, it will be necessary to highlight this fact and issue a request that follow-up pointed observations be made. This is particularly important in the case of objects which will benefit from observation with limited-lifespan space-based observatories, such as Herschel. Time awarded ------------ SASSy has been awarded 500 hours (the equivalent of approximately 42 nights) for the initial 2-year phase to map nearly 5000 square degrees of sky. If JCMT operations continue beyond the initial phase, then a further 1000 hours of time are available to continue SASSy. The survey will begin immediately after full commissioning of SCUBA-2, currently expected to be August 2009. Since SASSy will only be feasible with a full complement of four detector arrays, delays in their delivery or sub-standard operation will have a significant impact on the ability to perform this survey. Expected resources ------------------ The SASSy project will require access to significant computing resources. Raw data will amount to up to 500 GB per night, and will be processed in chunks of up to 64 GB at a time, to produce maps which are a few GB in size (depending on the final requested pixelisation, which is still to be determined). Processing with the SCUBA-2 pipeline requires a working space of approximately four times the raw data amount. Current estimates of the processing time suggest that 1 night of data (12 hours) requires 48 CPU hours (this assumes that for SASSy, data from the shorter wavelength array will be too noisy to reduce). The mapping strategy is to divide the sky into a series of tiles which are then mosaiced together into the final map. This can be considered as 2 distinct strips, each essentially 10 degrees wide by 100 degrees long (1 Galactic and 1 extragalactic). Initially the tiles will be processed independently, with consistency checks performed on the overlap regions. The above time estimate assumes that the processing is carried out loading all of the data into memory at once. The data will be processed in chunks of about 15 minutes (corresponding to an area of roughly 1.5 degrees x 1.5 degrees), which requires at least 50 GB of memory. With less memory, the processing time increases drastically (an order of magnitude) unless the map area being processed is also decreased proportionately. The reason for this is that the data reduction pipeline uses an iterative algorithm, and any data that are not available in memory must be read repeatedly from the hard disk. The minimum memory requirement for reasonable performance is 16 GB. Access to resources will be primarily driven by the data collection schedule. Requests for re-processing may be made at any time, but are most likely after new data collection. Most re-processing requests will require the same resources as the initial processing. However, in some cases it may be advantageous to process larger amounts of data as a single unit, namely for fields which have been observed multiple times or regions of complex source structure which would benefit from being treated as a single region. Such large regions cannot be processed entirely in memory and will require significant CPU time as described above. Currently we have no estimate for how often we would need to carry out re-processing of larger regions. As described above, a centrally-hosted platform for collaboration will be vital for the success of the project. A public-facing web page and private wiki would serve the project well. Major processing steps ---------------------- The data processing must be completed in near real-time to avoid a backlog forming. This is defined as 24 hours for 12 hours of data. The processing is automated with a pipeline tuned specifically for processing SASSy data, yielding images and source catalogues as basic data products. The data products must then be examined visually by one of the SASSy team for quality assessment. A request to re-process the data (with modified inputs to the pipeline) may then be submitted, or the products are marked as satisfactory. APPENDIX Science Use Cases for SASSy --------------------------- This appendix describes, in detail, some example tasks the scientists need to carry out to achieve the science goals of SASSy. 1. Example Science Use Case (SUC): Timely access to data -------------------------------------------------------- - Access to pipeline-processed images by next (local) morning after observation completed - Transfer to own machine - Access to raw data within 12 hours of observation completed - Transfer raw data to dedicated DR machine - A mechanism is needed for approval by multiple users 2. Example SUC: Monitor survey progress --------------------------------------- - Ability to determine list of completed observations - Show relation between survey field ID and JCMT data files - Tabulate completion statistics - Visual representation of completion statistics - Visual representation of progress - ability to plot completion on map of sky - ability to plot region of planned 2/5 year coverage - ability to plot/identify imminently observable fields - completion map should indicate number of times a field has been observed. 3. Example SUC: Data Quality Assessment of map tiles ---------------------------------------------------- - Access online assessment of observations - Ability to plot time variable quantities (tau, seeing, pointing) - Access offline assessment of observations - Raw data issues - Automatically flag suspect scans - Noise, NEFD performance - Largest believable scale in image - Ability to grade/reject observations based on survey criteria - Inspect images visually for integrity, consistency with surrounding fields - Data map - Noise map(s) - Comparison of features in signal map with noise map - Check of power spectra of maps - Need estimate of photometric (calibration) & astrometric accuracy - Ability to register feedback/notes/comments regarding data quality - Contact survey coordinators/data handlers - Notify survey mailing list of updates/new data - View/preserve history of actions 4. Example SUC: Data Quality Assessment of mosaicing ---------------------------------------------------- - Run mosaicing step on map tiles - need to be able to select tiles for inclusion - choice of weighting scheme - different pixelization schemes for regridding - Comparison of overlap areas - Visualisation tools for comparing overlap region between tiles and in mosaics themselves - Should include Healpix pixelization, with total view as strip in Galactic coordinates within Healpix - Possibility of rejecting tiles at this stage and requesting repeat observations - Ability to re-run map-making pipeline over areas larger by factor of 4 - Comparison of mosaicing results with smaller and larger tiles 5. Example SUC: Object/source detection --------------------------------------- - Visually inspect results from source extraction - Point/compact sources (SExtractor) - Extended sources (CUPID) - Ability to flag/exclude unreliable extractions from catalogue - Plot known objects from user-selectable catalogues of compact/extended sources - Construct spectral energy distributions for sources with multi-wavelength data - Ability to highlight variable sources - Ability to modify source extraction configuration parameters - Ability to accept/store modified configuration parameters - Ability to run new source extraction tasks - Assess need for dedicated pointed follow-up with SCUBA-2 - Sources with low signal-to-noise to be re-observed - Variable sources to be followed up - Bright point sources to be assessed for suitability as calibrators - View/preserve history of actions 6. Example SUC: Reprocessing/re-reduction ----------------------------------------- - Ability to request re-processing of raw data - Ability to flag/exclude individual scans - Ability to process completed fields as a single entity - Ability to group related fields and process the data as a single unit for optimum image reconstruction - Visual tool for grouping fields (click on fields) - Text/tabular tool to select relevant data - Ability to modify processing parameters - Ability to experiment with modified parameters - Data products marked experimental by user - Ability to run source extraction on reprocessed data - Ability to allow reprocessed data to be accepted as definitive - View/preserve history of actions