| Name : condor
| |
| Version : 6.6.6
| Vendor : (none)
|
| Release : lcg3_sl3
| Date : 2004-10-05 16:40:56
|
| Group : Condor
| Source RPM : condor-6.6.6-lcg3_sl3.src.rpm
|
| Size : 165.06 MB
| |
| Packager : (none)
| |
| Summary : condor 6.6.6 with lcg patches
|
Description :
Condor/Condor-G version 6.6.6
05 Oct 2004: This is RPM condor-6.6.6-lcg3 Similar to lcg2 but:
- changes in condor_gridmanager to:
+ Fix file descriptor leak + Add extra checking on the format of job contact strings + Changed default GRIDMANAGER_JM_EXIT_LIMIT limit to 30 seconds + Disabled the LCG patch to take (possibly) new contact information for the result of an attempt to restart an already running JM. Globus apparently sometimes returns some previous contact string and not the contact of the most recently running JM. + Changed gridmanager state transition from REGISTER->STDIO_UPDATE->SUBMIT_COMMIT->... to REGISTER->SUBMIT_COMMIT->STDIO_UPDATE->...
to avoid a problem where by a JM waiting for the submit commit signal neveral replies to the stdio register gram signal. + Upon receipt of a callback from a job indicating that the JM should be STOPPING, the stop timer is started (if it wasn\'t already) so that a subsequent restart that shows the old JM is still running will also be subjected to the GRIDMANAGER_JM_EXIT_LIMIT timer.
- added check on the state file age in the grid_monitor. (24 hours). State files for which the state hasn\'t changed in this time, and are currently in state DONE, FAILED or UNSUBMITTED will be removed from disc.
29 Sep 2004: This is RPM condor-6.6.6-lcg2 Similar to lcg1 but:
- change in gahp_server (version now reported as lcg2) to reacquire credential every time the default proxy is selected
- changed grid_monitor.sh to avoid calling the jobmanager poll() method for jobs previously recorded as DONE or FAILED
21 Sep 2004: This build contains some lcg patches and is known as version 6.6.6 (lcg1). In particular:
(+) The globus version linked in based on \'2.4.3.vdt.1.1.14\' with some patches. The globus patches include: - Some memory leak fixes - Additional checks in some gss_assist routes on token length - Timeouts on asynchronous reads (including secure connection establishment phase and in the gram protocol) - changes in the gass transfer https protocol with respect to the \'chunk footer\' delineation for zero length chunks.
(+) The gahp_server contains some replacement routines for private globus internals (to avoid some scaling problems with the \'ez\' gass server implimenation). The private routines included with the gahp_server were updated to reflect the current globus version.
The gahp_server handing of cached proxies is also patched, to prevent caching the credential structure itself, but only the proxy filename. The credential is reconstruted as necessary from the proxy file.
A small memory leak releated to setting an enviromment variable was also fixed, and some glibc malloc (ie mallopt) options set to try to reduce heap fragmentation.
(+) The condor_gridmanager includes several small patches:
- Change in the location that the condor gridmonitor agent sends its status and log files to. This was to recude the accumulation of many old log files seen on a busy resource broker.
- A \'resource\' manager object is not immediatly destroyed when the last job finished, but is kept for a period. (default 30 minutes, configurable with the paramater GRIDMANAGER_KEEP_EMPTY_RESOURCE_DELAY). This was to avoid unnesseary gridmonitor agents being started on remote CEs.
- During some failures, the gridmanager could repeatedly attempt to restart the jobmanager of a failed job. Changed the handling of the globus error code in the case of \'non restartable\' failures.
- Changed the interpreation of \'authentication failure\' when attempting to register a gram callback address with an already running jobmanager to assume that failure may be due to another user\'s jobmanager now running on the original contact port. An attempt is then made to restart the jobmanager.
- Attempt prevent the possibility of repeatly retying to stop and restart a JM, where the JM always returns \'jm already running\'. To do this a patch was added to the gridmanager to try to determine when a JM is found running when, accoring to previous actions, it should have already been stopped. If the gridmanager believes the JM should have been stopped more than GRIDMANAGER_JM_EXIT_LIMIT (default 10) seconds previously then the job is cleared, since it is assumed the JM is in a bad state that cannot be recovered.
- If a job manager is found already running during a restart attempt, the (possibly new) contact string is saved - and then the normal sequence is followed of registering the callback address.
- If the gridmonitor agent is temporarily abandoned at a site, for any reason, it is not desirable to awake all the jobmanagers. Therefore the requirement that the gridmonitor agent should be active was removed from those needed for a jobmanager to be shutdown.
- If an attempt is made to submitt a job to a site that has been marked as down in excess of a certain length of time, the job will be held with a reason of globus error 12 (error connecting). The default limit is 30 minutes, configurable via GRIDMANAGER_SUBMIT_TO_DOWN_RESOURCE_TIMEOUT.
(+) gridmonitor.sh an LCG version, based on an older condor gridmonitor.sh
|
RPM found in directory: /vol/rzm8/EGEE/gLite/APT/R3.0/rhel30.old/RPMS.externals |
Hmm ... It's impossible ;-) This RPM doesn't exist on any FTP server
Provides :
perl(Condor)
condor
Requires :